Pith. sign in

REVIEW 3 major objections 4 minor 2 cited by

Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper proposes that all human feedback for reward-based learning—preferences, ratings, demonstrations, corrections, gaze, and language—can be classified along nine dimensions and assessed by seven quality metrics, unifying…

desk verdict A genuinely useful taxonomy for RLHF feedback research, with an exhaustiveness claim that outruns the evidence and a scalar-reward formalization that needs to acknowledge its own information loss. read the letter →

arxiv 2411.11761 v2 pith:VAG54AWA submitted 2024-11-18 cs.LG cs.HC

classification cs.LGcs.HC
keywords reinforcementlearningfromhumanfeedbacktaxonomyinteractivemachinehuman-computerinteractionrewardqualitymetricshuman-in-the-loopmulti-type
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to answer two questions: what kinds of human feedback exist for training reinforcement learning agents, and what makes feedback good. It claims that every feedback type can be placed on nine dimensions—intent, expression form, engagement, target relation, content level, target actuality, temporal granularity, choice set size, and exclusivity—and that seven quality metrics (expressiveness, ease, definiteness, context independence, precision, unbiasedness, informativeness) govern how useful feedback is. If this is right, RLHF system builders gain a common map and vocabulary, and the field moves from preference-only feedback toward a richer design space of mixed feedback types.

What carries the argument

The central object is the nine-dimensional taxonomy (D1–D9) together with the formalization of a feedback instance as a mapping $F: \mathcal{T} \to r_{\text{fb}}$, produced by a translation algorithm $\phi$ that turns raw measurements into (target, value) pairs under a context encoding. The dimensions span three groups: human-centered (intent, expression form, engagement), interface-centered (target relation, content level, target actuality), and model-centered (temporal granularity, choice set size, exclusivity). The seven quality metrics (Q1–Q7) operationalize what makes feedback good from human, interface, and model perspectives, and the derived requirements (UI.R1–R4, FP.R1–R4, RM.R1–R3) connect the taxonomy to concrete system design.

What would settle it

Take a concrete corrective utterance like "Don't put the cup there, place it on the coaster, but only if the coaster is dry" and attempt to encode it as a single (target, scalar) pair under the paper's formalism; if the resulting scalar leaves the agent unable to distinguish the dry-coaster condition from the wet-coaster one, the scalar channel demonstrably loses information the taxonomy claims to cover. Alternatively, have independent annotators classify a held-out set of feedback utterances from the surveyed papers into the nine dimensions and measure agreement; low agreement would falsify the exhaustiveness and orthogonality claims.

Watch

Extended reading notes

Core claim

The central claim is that the space of human feedback for reward-based learning is structured: every feedback utterance, whether a thumbs-up, a preference between two replies, a gaze fixation, a physical correction, or a natural-language instruction, is a point in a nine-dimensional space. The paper formalizes feedback as a mapping $F: \mathcal{T} \to r_{\text{fb}}$ from a target set of trajectories to a scalar reward value, optionally conditioned on a context encoding derived from the feedback state, which decomposes into human, interface, and agent sub-states. On this foundation it builds the taxonomy, the seven quality metrics, and a set of requirements for the user interface, feedback processor, and reward model. The paper supports the claim by classifying 141 surveyed papers from 2008 to 2024 within the taxonomy.

Load-bearing premise

The framework assumes that every kind of human feedback can be compressed into a scalar reward value (plus optional context) without losing anything essential for learning the right behavior.

Editorial extensions

If this is right

  • RLHF systems can move beyond pairwise preferences: the framework licenses mixed feedback types, letting users choose the most natural channel (rate, correct, demonstrate, describe) at each moment.
  • Feedback quality becomes measurable along seven axes, enabling interface designers and reward-model trainers to compare feedback channels on the same terms.
  • Reward models must condition on context encodings $C$ to handle the fact that the same raw utterance means different things in different feedback states.
  • Querying strategies can be defined over the full nine-dimensional space, selecting not just which target to show but which feedback type to request.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The scalar-reward bottleneck is the framework's deepest assumption; a testable corollary is whether context-conditioned reward models that condition on the full feedback state actually recover information lost by scalar compression, something the paper motivates but does not demonstrate empirically.
  • The nine dimensions read naturally as an annotation schema; a natural next step the paper does not take is measuring inter-annotator agreement when independent coders classify feedback utterances, which would convert the exhaustiveness claim into a quantitative one.
  • The paper's own cited warning that 'humans are not Boltzmann distributions' cuts against the scalar formalism, so the framework is best read as a communication-surface map, with the reward-model semantics left as open work.
  • The quality metrics could be turned into a scoring rubric for RLHF interfaces, such as a checklist evaluating a system on Q1–Q7; the paper stops at defining the qualities.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a conceptual framework for human feedback in reward-based reinforcement learning. It introduces a taxonomy of feedback along nine dimensions (human-centered: intent, expression form, engagement; interface-centered: target relation, content level, target actuality; model-centered: temporal granularity, choice set size, exclusivity), along with seven quality metrics (expressiveness, ease, definiteness, context independence, precision, unbiasedness, informativeness). It formalizes feedback as a measurement m translated via a function phi into processed feedback F: T -> r_fb, optionally conditioned on a context encoding C, and derives system requirements for user interfaces, feedback processors, and reward models. The framework is supported by a survey of 141 papers classified in Appendix A and by an implemented prototype interface.

Significance. If taken as a design vocabulary rather than as a lossless formal characterization, the framework is a useful interdisciplinary contribution: it synthesizes a large and heterogeneous literature, connects HCI concerns (cognitive load, expressiveness) with ML concerns (reward-model learnability), and makes concrete, falsifiable design recommendations. The authors deserve credit for the extensive classification appendix, the explicit formalization that builds on the reward-rational choice framework [80], and the derived design requirements (UI.R1-R4, FP.R1-R4, RM.R1-R3), which are concrete enough to guide system builders. However, the paper's stronger claims—that the taxonomy is exhaustive and that diverse feedback types can be unified without loss—are not established by the evidence presented. The iterative methodology in Section 3.1 guarantees classifiability by construction, and the scalar-reward formalism in Section 3.2 assumes, rather than demonstrates, that language, corrections, and demonstrations can be faithfully rendered as (target, scalar value, context) triples. The framework's value as a design space is real; its value as a unifying formal model is overstated.

major comments (3)
  1. [§3.2] The central unifying claim depends on the assumption that any feedback measurement can be losslessly represented as a target T, a scalar reward value r_fb, and a context encoding C (Eq. "F : T → r_fb ∈ R" and "φ : m → (T, r_fb, C)"). This is not established. For instructive and corrective language feedback (e.g., [56, 116, 164]), the propositional content ("use a different cleaner", "do this instead") has no natural location in a scalar r_fb, and the paper provides no rule for constructing T from arbitrary linguistic or physical-correction measurements. Context encoding C is defined only from contextual measurements m_ctxt (Section 3.2, paragraph "Based on our definition, we may identify two types of measurement variables"), not from the intrinsic content of the utterance. Consequently, two semantically distinct messages can map to the same (T, r_fb, C) triple, so the formalization assumes lossless unification rather than demonstrating it. The paper itself cites [113] ("Humans are not Boltzmann Distributions") arguing that reducing human feedback to scalar reward is a misspecification, but this objection is not integrated into the formalism or used to bound the information loss.
  2. [§3.1 and Appendix A] The taxonomy's exhaustiveness is guaranteed by construction. The authors specify "Exhaustiveness: All surveyed papers describing types of human feedback for agent training must be classifiable with the given dimensions" and describe an iterative process in which dimensions were refined until the surveyed papers were classifiable. Appendix A's classification of 141 papers therefore re-demonstrates the design target rather than independently validating the taxonomy. The paper also claims that "the framework is also implemented as a software system and validated to handle multiple use cases" and that it was "validated ... in expert interviews," but no protocol, results, or inter-coder agreement information is provided for either validation. This makes it impossible to assess the reliability, completeness, or objectivity of the taxonomy as a descriptive tool.
  3. [§3.7 and Table 1] The classification of established feedback types uses grey and blue checkmarks to indicate that a feedback type can have different attributes across papers, but the formal definitions in Sections 3.4–3.6 are categorical (e.g., D4: |T|=1 vs. |T|>1; D8: r_fb ∈ {0,1,≻,≺} vs. N vs. R). The paper does not explain how a feedback type with multiple grey-checked attributes corresponds to a "well-defined point" in the nine-dimensional space, nor how the orthogonality requirement (Section 3.1) is preserved when attributes can span dimensions. This ambiguity weakens the claim that the taxonomy allows any feedback channel to sit at a well-defined point on D1–D9 and that the dimensions are mutually exclusive.
minor comments (4)
  1. [§3.2] There are several typos in the formal definitions: "s_i ∈, a_j ∈ A" should read "s_i ∈ S, a_j ∈ A"; "s_i ⊂ s_i ⊆ S" in the Formal Definition of Content Level is garbled; and "the identify function" should be "the identity function."
  2. [§3.1] The survey is described as covering papers from 2008 to 2024, but Appendix A and the reference list include earlier works (e.g., [75, 97, 132] from 1996–2005, [58] from 2003). Please clarify whether the 2008–2024 window applies only to the "final survey" keyword search and not to the earlier candidate sets.
  3. [§4.1] In the "Optimizing expressiveness" paragraph, the phrase "open-ended, implicit or multi-modal feedback D2 options" should be "...feedback options related to D2" to avoid implying that D2 itself is a set of options; similarly, "act proactively D3" reads awkwardly.
  4. [§5.2.2] The sentence "Human-computer/human-robot interaction presents a huge opportunity to create novel feedback interactions" is vague; consider specifying which interaction modalities are missing from current RLHF systems.

Circularity Check

3 steps flagged · score 4.0 of 10

Moderate circularity: the taxonomy's exhaustiveness is a design constraint re-demonstrated on the same corpus, and the scalar formalization builds the unification into its definitions; self-citations add mild load.

  1. fitted input called prediction [Section 3.1, Methodology and Process, and Appendix A (Tables 3-5)]
    "Exhaustiveness: All surveyed papers describing types of human feedback for agent training must be classifiable with the given dimensions. ... Based on the final survey, we decided on nine dimensions and seven quality criteria summarized from the surveyed literature."

    The nine dimensions were selected through an iterative survey-and-refinement loop until the surveyed papers were classifiable; the exhaustiveness requirement is a design constraint on the dimension set, not an independent outcome. Appendix A then classifies the same 141-paper corpus, so the 'exhaustive categorization' claim restates the selection criterion rather than testing it. Coverage of the surveyed corpus is therefore true by construction, and no evidence is provided that the dimensions cover the broader space of possible human feedback beyond the papers that shaped them.

  2. self definitional [Section 3.2, Formalization of Human Feedback]
    "We define a processed human feedback instance F as a mapping from a target T to a feedback value r_fb: F :T→ r_fb ∈ R. ... To generate processed feedback, we need to design a translation algorithm φ:m→(T ,r_fb)."

    The paper's unifying claim over diverse feedback types is achieved by defining processed feedback as a scalar-valued mapping from a target. Language, gaze, physical corrections, and other semantically rich inputs are then forced into the (target, scalar value) schema by fiat. The framework provides no argument that this projection is lossless, and it even cites [113], 'Humans are not Boltzmann Distributions', to question scalar modeling of humans without integrating that objection. Thus the breadth of the formalization is an artifact of its own definition rather than a derived or empirically supported result.

1 more flagged steps
  1. self citation load bearing [Section 3.2, Formalization of Human Feedback]
    "Based on our previous discussions, a robust translation algorithm/reward modeling approach should take the feedback state into context [113]."

    The context-conditioning of the translation algorithm, which is built into the formalization's context encoding C, is justified by citation to the authors' own prior position paper [113] rather than by an external or independently derived argument. This self-citation is present and mildly load-bearing for the design of the formalism, though the broader framework also rests on substantial external literature, so it does not by itself force the paper's central claims.

full rationale

The paper is a conceptual framework, not an empirical prediction paper, so most of its content is definitional and descriptive rather than a fitted-parameter-then-prediction chain. However, the central exhaustiveness claim does reduce to the methodology in part: the nine dimensions were refined until the surveyed papers were classifiable, and Appendix A re-classifies that same corpus, so the taxonomy's coverage is a restatement of the design constraint. The scalar formalization in Section 3.2 also builds the unification into the definition of processed feedback F : T -> r_fb, so semantically distinct feedback types are 'unified' by construction rather than by demonstrated lossless translation. The authors' own prior work [113] is cited as support for the context-conditioning and for the limitations of scalar human modeling, introducing a self-citation layer, but the survey corpus is external and the classification is transparent. Weighing these, the score is moderate: the framework has independent content and is not vacuous, but some of its load-bearing claims are true by construction or by self-citation rather than by independent validation.

Assumptions & free parameters 3 free parameters · 5 assumptions · 3 invented entities

The central claim rests entirely on hand-chosen organizing decisions: the 9 dimensions, the 7 qualities, and the scalar (target, reward) encoding. These are neither derived from first principles nor validated empirically; the paper's stated requirements guarantee the corpus is classifiable, so the appendix demonstration is partly circular. The only external anchor is the 141-paper corpus itself, which is only partially shown.

free parameters (3)
  • Taxonomy dimension set D1-D9 and per-dimension attribute sets = 9 dimensions with 2-4 attributes each
    Sections 3.3-3.6; the set was refined iteratively by four authors and a six-expert session until the surveyed corpus was classifiable (Section 3.1); the count and attribute choices are hand-made, not derived.
  • Quality criterion set Q1-Q7 = 7 criteria across human, interface, and model perspectives
    Section 4; dozens of literature terms in Table 2 are condensed into seven criteria by the authors; the grouping is a design choice.
  • Survey corpus composition and size = 141 papers, 2008-2024, keyword-restricted
    Section 3.1; the corpus bounds the exhaustiveness claim; the appendix displays fewer classifications than 141, and the filter is author-chosen.
assumptions (5)
  • domain assumption The nine dimensions and their attribute sets are exhaustive and orthogonal over the space of human feedback.
    Section 3.1 lists exhaustiveness and orthogonality as design requirements satisfied through author discussions; Appendix A applies the framework to the corpus but no independent test of the requirement is provided.
  • domain assumption The human-AI communication space decomposes into human, interface, and model actors with the stated goals of expressiveness, fidelity, and comprehensibility.
    Section 2.1; this framing is inherited from the authors' prior work [55] on communication processes in human-AI interaction and is assumed without independent justification.
  • ad hoc to paper All feedback types can be encoded as a measurement m translated into a processed feedback (T, r_fb) with a scalar value and an optional context encoding C.
    Section 3.2; the scalarization is the mechanism that lets the framework claim unification across feedback types, but losslessness is not proven or empirically tested.
  • domain assumption The reward-rational implicit choice framework [80] is a valid foundation for unifying feedback formalisms.
    Section 3.2 builds on Jeon et al. [80] and the authors' own RLHF-Blender [137]; the grounding-function concept is adopted without critical evaluation.
  • ad hoc to paper The expert interviews and implementation mentioned in Section 3.1 constitute validation of the framework.
    No protocol, participant count, or results are reported, so this claimed validation is not verifiable.
invented entities (3)
  • Nine-dimensional feedback taxonomy (D1-D9)
    purpose: Classify any human feedback type for reward-based learning.
    Core contribution; no external benchmark exists, and the dimensions were tuned until the authors' corpus was classifiable (Section 3.1).
  • Seven quality metrics (Q1-Q7)
    purpose: Assess feedback quality from human, interface, and model perspectives.
    Proposed measurement methods (response-time proxies, information gain, variance estimates) are not validated against human-subject data in this paper.
  • Feedback process formalization, including measurement m, target T, translation function phi, and feedback state f_s
    purpose: Supply a common notation connecting HCI and ML formalisms.
    Section 3.2 and Figure 4; descriptive notation with no theorems or predictions, and with errors such as D5's set-containment typo.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework." pith.science (2026). https://pith.science/paper/VAG54AWA

@misc{pith2026241111761,
  author       = {Pith},
  title        = {Pith review of: Mapping out the Space of Human Feedback for Reinforcement Learning: A Conceptual Framework},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VAG54AWA}},
  note         = {Machine review of arXiv:2411.11761}
}
read the original abstract

Reinforcement Learning from Human feedback (RLHF) has become a powerful tool to fine-tune or train agentic machine learning models. Similar to how humans interact in social contexts, we can use many types of feedback to communicate our preferences, intentions, and knowledge to an RL agent. However, applications of human feedback in RL are often limited in scope and disregard human factors. In this work, we bridge the gap between machine learning and human-computer interaction efforts by developing a shared understanding of human feedback in interactive learning scenarios. We first introduce a taxonomy of feedback types for reward-based learning from human feedback based on nine key dimensions. Our taxonomy allows for unifying human-centered, interface-centered, and model-centered aspects. In addition, we identify seven quality metrics of human feedback influencing both the human ability to express feedback and the agent's ability to learn from the feedback. Based on the feedback taxonomy and quality criteria, we derive requirements and design choices for systems learning from human feedback. We relate these requirements and design choices to existing work in interactive machine learning. In the process, we identify gaps in existing work and future research opportunities. We call for interdisciplinary collaboration to harness the full potential of reinforcement learning with data-driven co-adaptive modeling and varied interaction mechanics.

Figures

Figures reproduced from arXiv: 2411.11761 by the authors.

Figure 1
Figure 1. In the context of Reinforcement Learning, we present a conceptual framework of human feedback based on nine dimensions: We identify human-centered, interface-centered and model-centered dimension. For different types of feedback, we also need to consider human-centered, interface-centered and model-centered qualities which influence how feedback can be collected and processed. Reinforcement Learning from Human feedb… view at source ↗
Figure 2
Figure 2. Human-AI Interaction via Interactive Communication Interfaces: Modal Outputs are communicated [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 1
Figure 1. In this work, we focus on the first direction. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figures from the paper (13 more)
Figure 3
Figure 3. Figure 3: The methodology used to create the conceptual framework: During the [PITH_FULL_IMAGE:figures/full_fig_p006_3.png]
Figure 4
Figure 4. Figure 4: The feedback process formalized: Humans generate feedback for agents that they observe acting in [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The feedback state consists of three different sub-states: (1) The [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: In the context of Reinforcement Learning, we present a conceptual framework of human feedback based on nine dimensions: We identify human-centered, interface-centered and model-centered dimension. C, which can be utilized for reward model training, e.g., by training mu…
Figure 7
Figure 7. Figure 7: The basic system components for reward learning: A [PITH_FULL_IMAGE:figures/full_fig_p027_7.png]
Figure 8
Figure 8. Figure 8: A selection of proposed solutions to match human intentions – There have been several proposed solutions for user interfaces, the feedback processor, and the reward model. For user interfaces, we highlight (A) evaluative workflows, (B) demonstrations, and (C) descripti…
Figure 9
Figure 9. Figure 9: A Summary of Opportunities for Human-Centered Dimensions [PITH_FULL_IMAGE:figures/full_fig_p033_9.png]
Figure 10
Figure 10. Figure 10: A Summary of Opportunities for Interface-Centered Dimensions [PITH_FULL_IMAGE:figures/full_fig_p035_10.png]
Figure 11
Figure 11. Figure 11: A selection of proposed visualizations for different granularities in RL scenarios – several proposed solutions for user interfaces display elements of a reinforcement learning process at different granularities. 5.4.2 Choice Set Size D8 The available choice set is hi…
Figure 12
Figure 12. Figure 12: A Summary of Opportunities for Model-Centered Dimensions [PITH_FULL_IMAGE:figures/full_fig_p038_12.png]
Figure 13
Figure 13. Figure 13: Prototype implementation of an integrated system for multi-type human feedback in reinforcement [PITH_FULL_IMAGE:figures/full_fig_p039_13.png]
Figure 14
Figure 14. Figure 14: The publication years of surveyed publications. [PITH_FULL_IMAGE:figures/full_fig_p051_14.png]
Figure 15
Figure 15. Figure 15: Counts of surveyed papers for each attribute in Dimensions D1-D9 [PITH_FULL_IMAGE:figures/full_fig_p051_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Guided Reinforcement Learning: Addressing Training Bottlenecks through Policy Modulation

    cs.AI 2025-05 conditional novelty 6.0 of 10

    An LLM identifies critical states, suggests corrective actions, and assigns shaped rewards to refine an existing RL policy, beating several baselines in Pong and MuJoCo.

  2. Optimal Interactive Learning on the Job via Facility Location Planning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    COIL casts multi-task interactive robot learning as an uncapacitated facility location problem and uses approximation algorithms to plan skill, preference, and help queries that reduce human effort.

Reference graph

Works this paper leans on

230 extracted references · 26 canonical work pages · cited by 2 Pith papers

  1. [113]

    David Lindner and Mennatallah El-Assady. 2022. Humans are not Boltzmann Distributions: Challenges and Op- portunities for Modelling Human Feedback and Interaction in Reinforcement Learning. (June 2022). https: //doi.org/10.48550/arxiv.2206.13316 arXiv: 2206.13316

  2. [80]

    Hong Jun Jeon, Smitha Milli, and Anca D. Dragan. 2020. Reward-rational (implicit) choice: A unifying formalism for reward learning. http://arxiv.org/abs/2002.04833 arXiv:2002.04833 [cs]

  3. [1]

    Adadi and M

    A. Adadi and M. Berrada. 2018. Peeking Inside the Black-Box: A Survey on Explainable Artificial Intelligence (XAI). IEEE Access 6 (2018), 52138–52160

  4. [2]

    M Mehdi Afsar, Trafford Crump, and Behrouz Far. 2022. Reinforcement learning based recommender systems: A survey. Comput. Surveys 55, 7 (2022), 1–38

  5. [3]

    Sharath Chandra Akkaladevi, Matthias Plasch, Andreas Pichler, and Markus Ikeda. 2019. Towards Reinforcement based Learning of an Assembly Process for Human Robot Collaboration. Procedia Manufacturing 38 (2019), 1491–

  6. [4]

    Sharath Chandra Akkaladevi, Matthias Plasch, Andreas Pichler, and Markus Ikeda. 2019. Towards reinforcement based learning of an assembly process for human robot collaboration. Procedia Manufacturing 38 (2019), 1491–1498

  7. [5]

    Riad Akrour, Marc Schoenauer, and Michèle Sebag. 2012. April: Active preference learning-based reinforcement learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2012, Bristol, UK, September 24-28, 2012. Proceedings, Part II 23 . Springer, 116–131

  8. [6]

    Saleema Amershi, Maya Cakmak, William Bradley Knox, and Todd Kulesza. 2014. Power to the people: The role of humans in interactive machine learning. Ai Magazine 35, 4 (2014), 105–120

Show all 230 references
  1. [7]

    Saleema Amershi, James Fogarty, and Daniel Weld. 2012. ReGroup: Interactive Machine Learning for On-Demand Group Creation. Conference on Human Factors in Computing Systems - Proceedings (05 2012). https://doi.org/10.1145/ 2207676.2207680

  2. [8]

    Ofra Amir, Ece Kamar, Andrey Kolobov, and Barbara J. Grosz. 2016. Interactive Teaching Strategies for Agent Training. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence (New York, New York, USA) (IJCAI’16). AAAI Press, 804–811

  3. [9]

    Riku Arakawa, Sosuke Kobayashi, Yuya Unno, Yuta Tsuboi, and Shin-ichi Maeda. 2018. Dqn-tamer: Human-in-the-loop reinforcement learning with intractable feedback. arXiv preprint arXiv:1810.11748 (2018)

  4. [10]

    Reuben M Aronson and Henny Admoni. 2022. Gaze complements control input for goal prediction during assisted teleoperation. In Robotics science and systems

  5. [11]

    Christian Arzate Cruz and Takeo Igarashi. 2020. Mariomix: Creating aligned playstyles for bots with interactive reinforcement learning. In Extended abstracts of the 2020 annual symposium on computer-human interaction in play . 134–139

  6. [12]

    Christian Arzate Cruz and Takeo Igarashi. 2020. A survey on interactive reinforcement learning: Design principles and open challenges. DIS 2020 - Proceedings of the 2020 ACM Designing Interactive Systems Conference (July 2020), 1195–

  7. [13]

    Christian Arzate Cruz and Takeo Igarashi. 2020. A survey on interactive reinforcement learning: Design principles and open challenges. In Proceedings of the 2020 ACM designing interactive systems conference . 1195–1209

  8. [14]

    Brown, Jack Clark, Sam McCandlish, Chris Olah, and Jared Kaplan

    Amanda Askell, Yuntao Bai, Anna Chen, Dawn Drain, Deep Ganguli, Tom Henighan, Andy Jones, Nicholas Joseph, Benjamin Mann, Nova DasSarma, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Jackson Kernion, Kamal Ndousse, Catherine Olsson, Dario Amodei, Tom B. Brown, Jack Clark...

  9. [15]

    Akanksha Atrey, Kaleigh Clary, and David Jensen. 2020. Exploratory Not Explanatory: Counterfactual Analysis of Saliency Maps for Deep Reinforcement Learning. In International Conference on Learning Representations . https: //openreview.net/forum?id=rkl3m1BFDB , Vol. 1, No. 1, ...

  10. [16]

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al. 2022. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862 (2022)

  11. [17]

    Bowman, Zac Hatfield-Dodds, Ben Mann, Dario Amodei, Nicholas Joseph, Sam McCandlish, Tom Brown, and Jared Kaplan

    Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan P...

  12. [18]

    Chandrayee Basu, Erdem Bıyık, Zhixun He, Mukesh Singhal, and Dorsa Sadigh. 2019. Active learning of reward dynamics from hierarchical queries. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 120–127

  13. [19]

    Chandrayee Basu, Mukesh Singhal, and Anca D Dragan. 2018. Learning from richer human guidance: Augmenting comparison-based learning with feature queries. In Proceedings of the 2018 ACM/IEEE International Conference on Human-Robot Interaction. 132–140

  14. [20]

    M. G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling. 2013. The Arcade Learning Environment: An Evaluation Platform for General Agents. Journal of Artificial Intelligence Research 47 (jun 2013), 253–279

  15. [21]

    Jürgen Bernard, Marco Hutter, Matthias Zeppelzauer, Dieter Fellner, and Michael Sedlmair. 2017. Comparing visual- interactive labeling with active learning: An experimental study. IEEE transactions on visualization and computer graphics 24, 1 (2017), 298–308

  16. [22]

    Jürgen Bernard, Matthias Zeppelzauer, Markus Lehmann, Martin Müller, and Michael Sedlmair. 2018. Towards User-Centered Active Learning Algorithms. Computer Graphics Forum 37, 3 (June 2018), 121–132. https://doi.org/10. 1111/cgf.13406

  17. [23]

    Jürgen Bernard, Matthias Zeppelzauer, Michael Sedlmair, and Wolfgang Aigner. 2018. VIAL: a unified process for visual interactive labeling. The Visual Computer 34 (2018), 1189–1207

  18. [24]

    Adam Bignold, Francisco Cruz, Richard Dazeley, Peter Vamplew, and Cameron Foale. 2021. An Evaluation Methodology for Interactive Reinforcement Learning with Simulated Users. Biomimetics 6, 1 (2021). https://doi.org/10.3390/ biomimetics6010013

  19. [25]

    Adam Bignold, Francisco Cruz, Richard Dazeley, Peter Vamplew, and Cameron Foale. 2021. Persistent rule-based interactive reinforcement learning. Neural Computing and Applications (2021), 1–18

  20. [26]

    Adam Bignold, Francisco Cruz, Richard Dazeley, Peter Vamplew, and Cameron Foale. 2023. Human engagement providing evaluative and informative advice for interactive reinforcement learning.Neural Computing and Applications 35, 25 (2023), 18215–18230

  21. [27]

    Erdem Bıyık, Nicolas Huynh, Mykel J Kochenderfer, and Dorsa Sadigh. 2023. Active preference-based Gaussian process regression for reward learning and optimization. The International Journal of Robotics Research (2023), 02783649231208729

  22. [28]

    Andreea Bobu, Marius Wiggert, Claire Tomlin, and Anca D Dragan. 2021. Feature expansive reward learning: Rethinking human input. In Proceedings of the 2021 ACM/IEEE International Conference on Human-Robot Interaction . 216–224

  23. [29]

    Bradley Knox and Peter Stone

    W. Bradley Knox and Peter Stone. 2008. TAMER: Training an Agent Manually via Evaluative Reinforcement. In2008 7th IEEE International Conference on Development and Learning . 292–297. https://doi.org/10.1109/DEVLRN.2008.4640845

  24. [30]

    Satchuthananthavale RK Branavan, Harr Chen, Luke Zettlemoyer, and Regina Barzilay. 2009. Reinforcement learning for mapping instructions to actions. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natur...

  25. [31]

    Daniel S Brown, Yuchen Cui, and Scott Niekum. 2018. Risk-aware active inverse reinforcement learning. InConference on Robot Learning. PMLR, 362–372

  26. [32]

    Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum

    Daniel S. Brown, Wonjoon Goo, Prabhat Nagarajan, and Scott Niekum. 2019. Extrapolating Beyond Suboptimal Demon- strations via Inverse Reinforcement Learning from Observations. http://arxiv.org/abs/1904.06387 arXiv:1904.06387 [cs, stat]

  27. [33]

    Maya Cakmak, Siddhartha S Srinivasa, Min Kyung Lee, Jodi Forlizzi, and Sara Kiesler. 2011. Human preferences for robot-human hand-over configurations. In 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 1986–1993

  28. [34]

    Maya Cakmak and Andrea L. Thomaz. 2014. Eliciting good teaching from humans for machine learners. Artificial Intelligence 217 (2014), 198–215. https://doi.org/10.1016/j.artint.2014.08.005 , Vol. 1, No. 1, Article . Publication date: February 2025. 42 Y. Metz et al

  29. [35]

    Kate Candon, Jesse Chen, Yoony Kim, Zoe Hsu, Nathan Tsoi, and Marynel Vázquez. 2023. Nonverbal Human Signals Can Help Autonomous Agents Infer Human Preferences for Their Behavior. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems . 307–316

  30. [36]

    Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, et al. 2023. Open problems and fundamental limitations of reinforcement learning from human feedback. arXiv preprint a...

  31. [37]

    Colin Cherry. 1966. On human communication. (1966)

  32. [38]

    Mohamed Chetouani. 2021. Interactive Robot Learning: An Overview.ECCAI Advanced Course on Artificial Intelligence (2021), 140–172

  33. [39]

    Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, and Yoshua Bengio. 2018. BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning. 7th International Conference on Learning Representations, ICLR...

  34. [40]

    Vivienne Bihe Chi and Bertram F Malle. 2022. Instruct or evaluate: how people choose to teach norms to social robots. In 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI) . IEEE, 718–722

  35. [41]

    Bernard CK Choi and Anita WP Pak. 2004. A catalog of biases in questionnaires. Preventing chronic disease 2, 1 (2004), A13

  36. [42]

    Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep Reinforce- ment Learning from Human Preferences. 30 (2017), 4299–4307. https://proceedings.neurips.cc/paper/2017/file/ d5e2c0adad503c91f91df240d0cd4e49-Paper.pdf

  37. [43]

    Ho, Joseph L

    Yun-Shiuan Chuang, Xuezhou Zhang, Yuzhe Ma, Mark K. Ho, Joseph L. Austerweil, and Xiaojin Zhu. 2021. Using Ma- chine Teaching to Investigate Human Assumptions when Teaching Reinforcement Learners. arXiv:2009.02476 [cs.LG]

  38. [44]

    Laure Crochepierre, Lydia Boudjeloud-Assala, and Vincent Barbesant. 2022. Interactive reinforcement learning for symbolic regression from multi-format human-preference feedbacks. In31st International Joint Conference on Artificial Intelligence (IJCAI 2022)

  39. [45]

    Christian Arzate Cruz and Takeo Igarashi. 2021. Interactive explanations: Diagnosis and repair of reinforcement learning based agent behaviors. In 2021 IEEE Conference on Games (CoG) . IEEE, 01–08

  40. [46]

    Francisco Cruz, Adam Bignold, Hung Son Nguyen, Richard Dazeley, and Peter Vamplew. 2022. Broad-persistent Advice for Interactive Reinforcement Learning Scenarios. arXiv preprint arXiv:2210.05187 (2022)

  41. [47]

    Parisi, and Stefan Wermter

    Francisco Cruz, German I. Parisi, and Stefan Wermter. 2018. Multi-modal Feedback for Affordance-driven Interactive Reinforcement Learning. In 2018 International Joint Conference on Neural Networks (IJCNN) . 1–8. https://doi.org/10. 1109/IJCNN.2018.8489237

  42. [48]

    Yuchen Cui, Qiping Zhang, Brad Knox, Alessandro Allievi, Peter Stone, and Scott Niekum. 2021. The empathic framework for task learning from implicit human feedback. In Conference on Robot Learning . PMLR, 604–626

  43. [49]

    Felipe Leno Da Silva, Pablo Hernandez-Leal, Bilal Kartal, and Matthew E Taylor. 2020. Uncertainty-aware action advising for deep reinforcement learning agents. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 5792–5799

  44. [50]

    Joris De Winter, Albert De Beir, Ilias El Makrini, Greet Van de Perre, Ann Nowé, and Bram Vanderborght. 2019. Accelerating interactive reinforcement learning by human advice for an assembly task by a cobot. Robotics 8, 4 (2019), 104

  45. [51]

    Deshpande, Jeff Schneider, Deepak Pathak, David Held, and Benjamin Eysenbach

    Shuby V. Deshpande, Jeff Schneider, Deepak Pathak, David Held, and Benjamin Eysenbach. 2020.Towards Interpretable Reinforcement Learning, Interactive Visual- izations to Increase Insight . Master’s thesis. Carnegie Mellon University

  46. [52]

    Yan Duan, Xi Chen, Rein Houthooft, John Schulman, and Pieter Abbeel. 2016. Benchmarking Deep Reinforcement Learning for Continuous Control. arXiv:1604.06778 [cs.LG]

  47. [53]

    Dudley and Per Ola Kristensson

    John J. Dudley and Per Ola Kristensson. 2018. A Review of User Interface Design for Interactive Machine Learning. ACM Trans. Interact. Intell. Syst. 8, 2, Article 8 (jun 2018), 37 pages. https://doi.org/10.1145/3185517

  48. [54]

    Layla El Asri, Bilal Piot, Matthieu Geist, Romain Laroche, and Olivier Pietquin. 2016. Score-based Inverse Rein- forcement Learning. In Proceedings of the 2016 International Conference on Autonomous Agents & Multiagent Systems (Singapore, Singapore) (AAMAS ’16). International ...

  49. [55]

    Mennatallah El-Assady and Caterina Moruzzi. 2022. Which Biases and Reasoning Pitfalls Do Explanations Trigger? Decomposing Communication Processes in Human–AI Interaction. IEEE Computer Graphics and Applications 42, 6 (2022), 11–23

  50. [56]

    Ahmed Elgohary, Christopher Meek, Matthew Richardson, Adam Fourney, Gonzalo Ramos, and Ahmed Hassan Awadallah. 2021. NL-EDIT: Correcting Semantic Parse Errors through Natural Language Interaction. InProceedings of the 2021 Conference of the North American Chapter of the Associ...

  51. [57]

    Anestis Fachantidis, Matthew E Taylor, and Ioannis Vlahavas. 2017. Learning to teach reinforcement learning agents. Machine Learning and Knowledge Extraction 1, 1 (2017), 21–42

  52. [58]

    Jerry Alan Fails and Dan R. Olsen. 2003. Interactive Machine Learning. In Proceedings of the 8th International Conference on Intelligent User Interfaces (Miami, Florida, USA) (IUI ’03). Association for Computing Machinery, New York, NY, USA, 39–45. https://doi.org/10.1145/6040...

  53. [59]

    Taylor A Kessler Faulkner, Elaine Schaertl Short, and Andrea L Thomaz. 2020. Interactive reinforcement learning with inaccurate feedback. In 2020 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 7498–7504

  54. [60]

    Xuening Feng, Zhaohui JIANG, Timo Kaufmann, Eyke Hüllermeier, Paul Weng, and Yifei Zhu. 2024. Comparing Comparisons: Informative and Easy Human Feedback with Distinguishability Queries. In ICML 2024 Workshop on Models of Human Feedback for AI Alignment . https://openreview.net...

  55. [61]

    Patrick Fernandes, Aman Madaan, Emmy Liu, António Farinhas, Pedro Henrique Martins, Amanda Bertsch, José G. C. de Souza, Shuyan Zhou, Tongshuang Wu, Graham Neubig, and André F. T. Martins. 2023. Bridging the Gap: A Survey on Integrating (Human) Feedback for Natural Language Ge...

  56. [62]

    Pierre W Ferrez and José del R Millán. 2005. You are wrong!—automatic detection of interaction errors from brain waves. In Proceedings of the 19th international joint conference on artificial intelligence

  57. [63]

    Tesca Fitzgerald, Pallavi Koppol, Patrick Callaghan, Russell Quinlan Jun Hei Wong, Reid Simmons, Oliver Kroemer, and Henny Admoni. 2023. INQUIRE: INteractive querying for user-aware informative REasoning. In Conference on Robot Learning. PMLR, 2241–2250

  58. [64]

    Spencer Frazier and Mark Riedl. 2019. Improving deep reinforcement learning in minecraft with action advice. In Proceedings of the AAAI conference on artificial intelligence and interactive digital entertainment , Vol. 15. 146–152

  59. [65]

    Rachel Freedman, Rohin Shah, and Anca D. Dragan. 2021. Choice Set Misspecification in Reward Inference. CoRR abs/2101.07691 (2021). arXiv:2101.07691 https://arxiv.org/abs/2101.07691

  60. [66]

    Deep Ganguli, Liane Lovitt, Jackson Kernion, Amanda Askell, Yuntao Bai, and Saurav Kadavath et al. 2022. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv:2209.07858 [cs.CL]

  61. [67]

    Yang Gao, Christian M Meyer, and Iryna Gurevych. 2018. APRIL: Interactively learning to summarise by combining active preference learning and reinforcement learning. arXiv preprint arXiv:1808.09658 (2018)

  62. [68]

    Gaurav R Ghosal, Matthew Zurek, Daniel S Brown, and Anca D Dragan. 2022. The Effect of Modeling Human Rationality Level on Learning Rewards from Multiple Feedback Types. arXiv preprint arXiv:2208.10687 (2022)

  63. [69]

    Shane Griffith, Kaushik Subramanian, Jonathan Scholz, Charles L Isbell, and Andrea L Thomaz. 2013. Policy shaping: Integrating human feedback with reinforcement learning. Advances in neural information processing systems 26 (2013)

  64. [70]

    Lin Guan, Mudit Verma, Suna Sihang Guo, Ruohan Zhang, and Subbarao Kambhampati. 2021. Widening the pipeline in human-guided reinforcement learning with explanation and context-aware data augmentation. Advances in Neural Information Processing Systems 34 (2021), 21885–21897

  65. [71]

    David Gunning, Mark Stefik, Jaesik Choi, Timothy Miller, Simone Stumpf, and Guang-Zhong Yang. 2019. XAI—Explainable artificial intelligence. Science robotics 4, 37 (2019), eaay7120

  66. [72]

    Dylan Hadfield-Menell, Stuart J Russell, Pieter Abbeel, and Anca Dragan. 2016. Cooperative inverse reinforcement learning. Advances in neural information processing systems 29 (2016)

  67. [73]

    Daniel Harnack, Julie Pivin-Bachler, and Nicolás Navarro-Guerrero. 2023. Quantifying the effect of feedback frequency in interactive reinforcement learning for robotic tasks. Neural Computing and Applications 35, 23 (2023), 16931–16943

  68. [74]

    Brent Harrison, Upol Ehsan, and Mark O Riedl. 2017. Guiding reinforcement learning exploration using natural language. arXiv preprint arXiv:1707.08616 (2017)

  69. [75]

    Frederick Hayes-Roth, Philip Klahr, and David J Mostow. 2013. Advice taking and knowledge refinement: An iterative view of skill acquisition. In Cognitive skills and their acquisition . Psychology Press, 231–253

  70. [76]

    Tom Hosking, Phil Blunsom, and Max Bartolo. 2024. Human Feedback is not Gold Standard. InThe Twelfth International Conference on Learning Representations . https://openreview.net/forum?id=7W3GLNImfS

  71. [77]

    Eric Hsiung, Eric Rosen, Vivienne Bihe Chi, and Bertram F Malle. 2022. Learning reward functions from a combination of demonstration and evaluative feedback. In2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 807–811

  72. [78]

    Ahmed Hussein, Mohamed Medhat Gaber, Eyad Elyan, and Chrisina Jayne. 2017. Imitation Learning: A Survey of Learning Methods. 50, 2, Article 21 (apr 2017), 35 pages. https://doi.org/10.1145/3054912

  73. [79]

    Borja Ibarz, Jan Leike, Tobias Pohlen, Geoffrey Irving, Shane Legg, and Dario Amodei. 2018. Reward learning from human preferences and demonstrations in atari. Advances in neural information processing systems 31 (2018). , Vol. 1, No. 1, Article . Publication date: February 20...

  74. [81]

    Kaixuan Ji, Jiafan He, and Quanquan Gu. 2024. Reinforcement Learning from Human Feedback with Active Queries. arXiv:2402.09401 [cs.LG] https://arxiv.org/abs/2402.09401

  75. [82]

    Kshitij Judah, Saikat Roy, Alan Fern, and Thomas Dietterich. 2010. Reinforcement learning via practice and critique advice. In Proceedings of the AAAI conference on artificial intelligence , Vol. 24. 481–486

  76. [83]

    Arthur Juliani, Vincent-Pierre Berges, Ervin Teng, Andrew Cohen, Jonathan Harper, Chris Elion, Chris Goy, Yuan Gao, Hunter Henry, Marwan Mattar, and Danny Lange. 2020. Unity: A general platform for intelligent agents. arXiv preprint arXiv:1809.02627 (2020)

  77. [84]

    Timo Kaufmann, Paul Weng, Viktor Bengs, and Eyke Hüllermeier. 2023. A survey of reinforcement learning from human feedback. arXiv preprint arXiv:2312.14925 (2023)

  78. [85]

    Su Kyoung Kim, Elsa Andrea Kirchner, Arne Stefes, and Frank Kirchner. 2017. Intrinsic interactive reinforcement learning–Using error-related potentials for real world human-robot interaction. Scientific reports 7, 1 (2017), 1–16

  79. [86]

    W Bradley Knox and Peter Stone. 2009. Interactively shaping agents via human reinforcement: The TAMER framework. In Proceedings of the fifth international conference on Knowledge capture . 9–16

  80. [87]

    W Bradley Knox and Peter Stone. 2010. Combining manual feedback with subsequent MDP reward signals for reinforcement learning.. In AAMAS. 5–12

  81. [88]

    W Bradley Knox and Peter Stone. 2012. Reinforcement learning from human reward: Discounting in episodic tasks. In 2012 IEEE RO-MAN: The 21st IEEE international symposium on robot and human interactive communication . IEEE, 878–885

  82. [89]

    Jens Kober, J Andrew Bagnell, and Jan Peters. 2013. Reinforcement learning in robotics: A survey. The International Journal of Robotics Research 32, 11 (2013), 1238–1274

  83. [90]

    Dorothea Koert, Maximilian Kircher, Vildan Salikutluk, Carlo D’Eramo, and Jan Peters. 2020. Multi-channel interactive reinforcement learning for sequential tasks. Frontiers in Robotics and AI 7 (2020), 97

  84. [91]

    Pallavi Koppol. 2023. Interactive Machine Learning from Humans: Knowledge Sharing via Mutual Feedback . Ph. D. Dissertation. Carnegie Mellon University

  85. [92]

    Thomas Kosch, Jakob Karolus, Johannes Zagermann, Harald Reiterer, Albrecht Schmidt, and Paweł W Woźniak. 2023. A survey on measuring cognitive workload in human-computer interaction. Comput. Surveys 55, 13s (2023), 1–39

  86. [93]

    Samantha Krening. 2018. Newtonian action advice: Integrating human verbal instruction with reinforcement learning. arXiv preprint arXiv:1804.05821 (2018)

  87. [94]

    Samantha Krening and Karen M Feigh. 2018. Interaction algorithm effect on human experience with reinforcement learning. ACM Transactions on Human-Robot Interaction (THRI) 7, 2 (2018), 1–22

  88. [95]

    Samantha Krening, Brent Harrison, Karen M Feigh, Charles Lee Isbell, Mark Riedl, and Andrea Thomaz. 2016. Learning from explanations using sentiment and advice in RL. IEEE Transactions on Cognitive and Developmental Systems 9, 1 (2016), 44–55

  89. [96]

    Julia Kreutzer, Shahram Khadivi, Evgeny Matusov, and Stefan Riezler. 2018. Can Neural Machine Translation be Improved with User Feedback?. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn...

  90. [97]

    Gregory Kuhlmann, Peter Stone, Raymond Mooney, and Jude Shavlik. 2004. Guiding a reinforcement learner with natural language advice: Initial results in RoboCup soccer. In The AAAI-2004 workshop on supervisory control of learning and adaptive systems . San Jose, CA

  91. [99]

    Minae Kwon, Siddharth Karamcheti, Mariano-Florentino Cuellar, and Dorsa Sadigh. 2021. Targeted data acquisition for evolving negotiation agents. In International Conference on Machine Learning . PMLR, 5894–5904

  92. [100]

    Cassidy Laidlaw and Stuart Russell. 2021. Uncertain decisions facilitate better preference learning. Advances in Neural Information Processing Systems 34 (2021), 15070–15083

  93. [101]

    Matthew V Law, Zhilong Li, Amit Rajesh, Nikhil Dhawan, Amritansh Kwatra, and Guy Hoffman. 2021. Hammers for Robots: Designing Tools for Reinforcement Learning Agents. In Designing Interactive Systems Conference 2021 . 1638–1653

  94. [102]

    Kimin Lee, Laura Smith, Anca Dragan, and Pieter Abbeel. 2021. B-pref: Benchmarking preference-based reinforcement learning. arXiv preprint arXiv:2111.03026 (2021). , Vol. 1, No. 1, Article . Publication date: February 2025. Mapping out the Space of Human Feedback for Reinforce...

  95. [103]

    Guangliang Li, Hamdi Dibeklioğlu, Shimon Whiteson, and Hayley Hung. 2020. Facial feedback for reinforcement learning: a case study and offline analysis using the TAMER framework. Autonomous Agents and Multi-Agent Systems 34 (2020), 1–29

  96. [104]

    Guangliang Li, Randy Gomez, Keisuke Nakamura, and Bo He. 2019. Human-Centered Reinforcement Learning: A Survey. IEEE Transactions on Human-Machine Systems 49, 4 (Aug. 2019), 337–349. https://doi.org/10.1109/THMS. 2019.2912447 Publisher: Institute of Electrical and Electronics ...

  97. [105]

    Guangliang Li, Bo He, Randy Gomez, and Keisuke Nakamura. 2018. Interactive reinforcement learning from demonstration and human evaluative feedback. In 2018 27th IEEE International Symposium on Robot and Human Interactive Communication (RO-MAN). IEEE, 1156–1162

  98. [106]

    Guangliang Li, Hayley Hung, Shimon Whiteson, and W Bradley Knox. 2013. Using informative behavior to increase engagement in the tamer framework. In Proceedings of the 2013 international conference on Autonomous agents and multi-agent systems. 909–916

  99. [107]

    Kejun Li, Maegan Tucker, Erdem Bıyık, Ellen Novoseller, Joel W Burdick, Yanan Sui, Dorsa Sadigh, Yisong Yue, and Aaron D Ames. 2021. Roial: Region of interest active learning for characterizing exoskeleton gait preference landscapes. In 2021 IEEE International Conference on Ro...

  100. [108]

    Mengxi Li, Alper Canberk, Dylan P Losey, and Dorsa Sadigh. 2021. Learning human objectives from sequences of physical corrections. In 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 2877–2883

  101. [109]

    Cristea, and Yunzhan Zhou

    Zhaoxing Li, Lei Shi, Alexandra I. Cristea, and Yunzhan Zhou. 2021. A Survey of Collaborative Reinforcement Learning: Interactive Methods and Design Patterns. In Designing Interactive Systems Conference 2021 (Virtual Event, USA) (DIS ’21). Association for Computing Machinery, ...

  102. [110]

    Jessy Lin, Daniel Fried, Dan Klein, and Anca Dragan. 2022. Inferring rewards from language in context. arXiv preprint arXiv:2204.02515 (2022)

  103. [111]

    Jinying Lin, Zhen Ma, Randy Gomez, Keisuke Nakamura, Bo He, and Guangliang Li. 2020. A review on interactive reinforcement learning from human social feedback. IEEE Access 8 (2020), 120757–120765

  104. [112]

    Zhiyu Lin, Brent Harrison, Aaron Keech, and Mark O Riedl. 2017. Explore, exploit or listen: Combining human feedback and policy model to speed up deep reinforcement learning in 3d worlds. arXiv preprint arXiv:1709.03969 (2017)

  105. [114]

    David Lindner, Rohin Shah, Pieter Abbeel, and Anca Dragan. 2021. Learning what to do by simulating the past. arXiv preprint arXiv:2104.03946 (2021)

  106. [115]

    Bing Liu, Gokhan Tür, Dilek Hakkani-Tür, Pararth Shah, and Larry Heck. 2018. Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems. InProceedings of the 2018 Conference of the North American Chapter of the Association for Com...

  107. [116]

    Huihan Liu, Alice Chen, Yuke Zhu, Adith Swaminathan, Andrey Kolobov, and Ching-An Cheng. 2023. Interactive Robot Learning from Verbal Correction. arXiv preprint arXiv:2310.17555 (2023)

  108. [117]

    Huihan Liu, Soroush Nasiriany, Lance Zhang, Zhiyao Bao, and Yuke Zhu. 2022. Robot Learning on the Job: Human- in-the-Loop Autonomy and Learning During Deployment. arXiv preprint arXiv:2211.08416 (2022)

  109. [118]

    YuXuan Liu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. 2018. Imitation from observation: Learning to imitate behaviors from raw video via context translation. In 2018 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 1118–1125

  110. [120]

    Robert Loftin, Bei Peng, James MacGlashan, Michael L Littman, Matthew E Taylor, Jeff Huang, and David L Roberts

  111. [121]

    Zvezdan Lončarević, Aleš Ude, Bojan Nemec, Andrej Gams, et al. 2018. User feedback in latent space robotic skill learning. In 2018 IEEE-RAS 18th International Conference on Humanoid Robots (Humanoids) . IEEE, 270–276

  112. [122]

    Manuel Lopes, Thomas Cederbourg, and Pierre-Yves Oudeyer. 2011. Simultaneous acquisition of task and feedback models. In 2011 IEEE International Conference on Development and Learning (ICDL) , Vol. 2. 1–7. https://doi.org/10. 1109/DEVLRN.2011.6037359

  113. [123]

    Dylan P Losey, Andrea Bajcsy, Marcia K O’Malley, and Anca D Dragan. 2022. Physical interaction as communication: Learning robot objectives online from human corrections. The International Journal of Robotics Research 41, 1 (2022), , Vol. 1, No. 1, Article . Publication date: F...

  114. [124]

    Dylan P Losey and Marcia K O’Malley. 2018. Including uncertainty when learning from human corrections. In Conference on Robot Learning . PMLR, 123–132

  115. [125]

    Fogarty, and Yang Li

    Hao Lü, James A. Fogarty, and Yang Li. 2014. Gesture Script: Recognizing Gestures and Their Structure Using Rendering Scripts and Interactively Trained Parts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (Toronto, Ontario, Canada) (CHI ’14). As...

  116. [126]

    Jelena Luketina, Nantas Nardelli, Gregory Farquhar, Jakob Foerster, Jacob Andreas, Edward Grefenstette, Shimon Whiteson, and Tim Rocktäschel. 2019. A Survey of Reinforcement Learning Informed by Natural Language. In Proceedings of the Twenty-Eighth International Joint Conferen...

  117. [127]

    Who’sa Good Robot?!

    P Lukowicz et al. 2023. “Who’sa Good Robot?!” Designing Human-Robot Teaching Interactions Inspired by Dog Training. In HHAI 2023: Augmenting Human Intellect: Proceedings of the Second International Conference on Hybrid Human-Artificial Intelligence, Vol. 368. IOS Press, 310

  118. [128]

    Jieliang Luo, Sam Green, Peter Feghali, George Legrady, C¸Etin, and Kaya Koç. 2018. Visual Diagnostics for Deep Reinforcement Learning Policy Development. (Sept. 2018). https://arxiv.org/abs/1809.06781v2 arXiv: 1809.06781 ISBN: 1809.06781v2

  119. [129]

    Daoming Lyu, Fangkai Yang, Bo Liu, and Steven Gustafson. 2019. SDRL: interpretable and data-efficient deep reinforcement learning leveraging symbolic planning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 2970–2977

  120. [130]

    James MacGlashan, Mark K Ho, Robert Loftin, Bei Peng, Guan Wang, David L Roberts, Matthew E Taylor, and Michael L Littman. 2017. Interactive learning from policy-dependent human feedback. In International conference on machine learning. PMLR, 2285–2294

  121. [131]

    Richard Maclin, Jude Shavlik, Lisa Torrey, Trevor Walker, and Edward Wild. 2005. Giving advice about preferred actions to reinforcement learners via knowledge-based kernel regression. In AAAI. 819–824

  122. [132]

    Richard Maclin and Jude W Shavlik. 1996. Creating advice-taking reinforcement learners. Machine Learning 22 (1996), 251–281

  123. [133]

    Dietterich, Rachel Houtman, Claire Montgomery, and Ronald Metoyer

    Sean McGregor, Hailey Buckingham, Thomas G. Dietterich, Rachel Houtman, Claire Montgomery, and Ronald Metoyer. 2015. Facilitating testing and debugging of Markov Decision Processes with interactive visualization. In 2015 IEEE Symposium on Visual Languages and Human-Centric Com...

  124. [134]

    Shaunak A Mehta and Dylan P Losey. 2022. Unified learning from demonstrations, corrections, and preferences during physical human-robot interaction. arXiv preprint arXiv:2207.03395 (2022)

  125. [135]

    Shaunak A Mehta, Forrest Meng, Andrea Bajcsy, and Dylan P Losey. 2024. StROL: Stabilized and Robust Online Learning from Humans. IEEE Robotics and Automation Letters (2024)

  126. [136]

    Marcel Menner, Lukas Neuner, Lars Lünenburger, and Melanie N Zeilinger. 2020. Using human ratings for feedback control: A supervised learning approach with application to rehabilitation robotics. IEEE Transactions on Robotics 36, 3 (2020), 789–801

  127. [137]

    Yannick Metz, David Lindner, Raphaël Baur, Daniel Keim, and Mennatallah El-Assady. 2023. RLHF-Blender: A Configurable Interactive Interface for Learning from Diverse Human Feedback. arXiv preprint arXiv:2308.04332 (2023)

  128. [138]

    Cristian Millán, Bruno JT Fernandes, and Francisco Cruz. 2019. Human feedback in continuous actor-critic reinforce- ment learning.. In ESANN

  129. [139]

    Cristian C Millan-Arias, Bruno JT Fernandes, Francisco Cruz, Richard Dazeley, and Sergio Fernandes. 2021. A robust approach for continuous interactive actor-critic algorithms. IEEE Access 9 (2021), 104242–104260

  130. [140]

    Sören Mindermann, Rohin Shah, Adam Gleave, and Dylan Hadfield-Menell. 2018. Active inverse reward design.arXiv preprint arXiv:1809.03060 (2018)

  131. [141]

    Ithan Moreira, Javier Rivas, Francisco Cruz, Richard Dazeley, Angel Ayala, and Bruno Fernandes. 2020. Deep reinforcement learning with interactive feedback in a human–robot environment. Applied Sciences 10, 16 (2020), 5574

  132. [142]

    Vivek Myers, Erdem Biyik, Nima Anari, and Dorsa Sadigh. 2022. Learning multimodal rewards from rankings. In Conference on Robot Learning . PMLR, 342–352

  133. [143]

    Anis Najar and Mohamed Chetouani. 2021. Reinforcement learning with human advice: a survey. Frontiers in Robotics and AI 8 (2021), 584075

  134. [144]

    Benjamin A Newman, Christopher Jason Paxton, Kris Kitani, and Henny Admoni. 2023. Towards Online Adaptation for Autonomous Household Assistants. In Companion of the 2023 ACM/IEEE International Conference on Human-Robot Interaction. 506–510. , Vol. 1, No. 1, Article . Publicati...

  135. [145]

    Andrew Y Ng, Daishi Harada, and Stuart Russell. 1999. Policy invariance under reward transformations: Theory and application to reward shaping. In Icml, Vol. 99. Citeseer, 278–287

  136. [146]

    Andrew Y Ng, Stuart Russell, et al. 2000. Algorithms for inverse reinforcement learning.. In Icml, Vol. 1. 2

  137. [147]

    Phillip Odom and Sriraam Natarajan. 2015. Active advice seeking for inverse reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 29

  138. [148]

    Takato Okudo and Seiji Yamada. 2021. Subgoal-based reward shaping to improve efficiency in reinforcement learning. IEEE Access 9 (2021), 97557–97568

  139. [149]

    Fredrik Olsson. 2009. A literature survey of active machine learning in the context of natural language processing. (2009)

  140. [150]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang Sandhini Agar- wal Katarina Slama Alex Ray John Schulman Jacob Hilton Fraser Kelton Luke Miller Maddie Simens Amanda Askell, Peter Welinder Paul Christiano, Jan Leike, and Ryan Low...

  141. [151]

    Rok Pahič, Zvezdan Lončarević, Aleš Ude, Bojan Nemec, and Andrej Gams. 2018. User Feedback in Latent Space Robotic Skill Learning. In 2018 IEEE-RAS 18th International Conference on Humanoid Robots (Humanoids) . 270–276. https://doi.org/10.1109/HUMANOIDS.2018.8624972

  142. [152]

    Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn. 2024. Disentangling length from quality in direct preference optimization. arXiv preprint arXiv:2403.19159 (2024)

  143. [153]

    Miriam Punzi, Nicolas Ladeveze, Huyen Nguyen, and Brian Ravenet. 2022. ImCasting: Nonverbal Behaviour Re- inforcement Learning of Virtual Humans through Adaptive Immersive Game. In 27th International Conference on Intelligent User Interfaces. 62–65

  144. [154]

    Nikaash Puri, Sukriti Verma, Piyush Gupta, Dhruv Kayastha, Shripad Deshmukh, Balaji Krishnamurthy, and Sameer Singh. 2020. Explain Your Move: Understanding Agent Actions Using Specific and Relevant Feature Attribution. In Intl. Conf. on Learning Representations . https://openr...

  145. [155]

    Syed Ali Raza, Benjamin Johnston, and Mary-Anne Williams. 2016. Reward from demonstration in interactive reinforcement learning. In The Twenty-Ninth International Flairs Conference

  146. [156]

    Syed Ali Raza and Mary-Anne Williams. 2020. Human feedback as action assignment in interactive reinforcement learning. ACM Transactions on Autonomous and Adaptive Systems (TAAS) 14, 4 (2020), 1–24

  147. [157]

    Keita Saito, Akifumi Wachi, Koki Wataoka, and Youhei Akimoto. 2023. Verbosity Bias in Preference Labeling by Large Language Models. arXiv:2310.10076 [cs.CL] https://arxiv.org/abs/2310.10076

  148. [158]

    Lisa Scherf, Cigdem Turan, and Dorothea Koert. 2022. Learning from Unreliable Human Action Advice in Interactive Reinforcement Learning. In 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids) . IEEE, 895–902

  149. [159]

    Mariah L Schrum, Erin Hedlund-Botti, Nina Moorman, and Matthew C Gombolay. 2022. Mind meld: Personalized meta-learning for robot-centric imitation learning. In 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 157–165

  150. [160]

    Burr Settles. 2009. Active learning literature survey. (2009)

  151. [161]

    Rita Sevastjanova, Fabian Beck, Basil Ell, Cagatay Turkay, Rafael Henkin, Miriam Butt, Daniel A Keim, and Mennatallah El-Assady. 2018. Going beyond visualization: Verbalization as complementary medium to explain machine learning models. In Workshop on Visualization for AI Expl...

  152. [162]

    Rita Sevastjanova, Wolfgang Jentner, Fabian Sperrle, Rebecca Kehlbeck, Jürgen Bernard, and Mennatallah El-assady

  153. [163]

    Rohin Shah, Pedro Freire, Neel Alex, Rachel Freedman, Dmitrii Krasheninnikov, Lawrence Chan, Michael D Dennis, Pieter Abbeel, Anca Dragan, and Stuart Russell. 2020. Benefits of assistance over reward learning. (2020)

  154. [164]

    Pratyusha Sharma, Balakumar Sundaralingam, Valts Blukis, Chris Paxton, Tucker Hermans, Antonio Torralba, Jacob Andreas, and Dieter Fox. 2022. Correcting robot plans with natural language feedback.arXiv preprint arXiv:2204.05186 (2022)

  155. [165]

    Isaac Sheidlower, Allison Moore, and Elaine Short. 2022. Keeping Humans in the Loop: Teaching via Feedback in Continuous Action Space Environments. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 863–870

  156. [166]

    Alvin Shek, Bo Ying Su, Rui Chen, and Changliu Liu. 2023. Learning from physical human feedback: An object- centric one-shot adaptation method. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 9910–9916

  157. [167]

    Tan, and Patrice Simard

    Michael Shilman, Desney S. Tan, and Patrice Simard. 2006. CueTIP: A Mixed-Initiative Interface for Correcting Handwriting Errors. In Proceedings of the 19th Annual ACM Symposium on User Interface Software and Technology , Vol. 1, No. 1, Article . Publication date: February 202...

  158. [168]

    Hilson Shrestha, Kathleen Cachel, Mallak Alkhathlan, Elke Rundensteiner, and Lane Harrison. 2022. FairFuse: Interactive Visual Support for Fair Consensus Ranking. In 2022 IEEE Visualization and Visual Analytics (VIS) . IEEE, 65–69

  159. [169]

    ACM Transactions on Interactive Intelligent Systems 11, 3-4 (Dec

    QuestionComb: A Gamification Approach for the Visual Explanation of Linguistic Phenomena through Interactive Labeling. ACM Transactions on Interactive Intelligent Systems 11, 3-4 (Dec. 2021), 1–38. https://doi.org/10. 1145/3429448

  160. [170]

    Shivam Singhal, Cassidy Laidlaw, and Anca Dragan. 2024. Scalable Oversight by Accounting for Unreliable Feedback. In ICML 2024 Workshop on Models of Human Feedback for AI Alignment. https://openreview.net/forum?id=Noy5wbyiCS

  161. [171]

    Anand Siththaranjan, Cassidy Laidlaw, and Dylan Hadfield-Menell. 2023. Distributional Preference Learning: Under- standing and Accounting for Hidden Context in RLHF. arXiv preprint arXiv:2312.08358 (2023)

  162. [172]

    Patrikk D Sørensen, Jeppeh M Olsen, and Sebastian Risi. 2016. Breeding a diversity of super mario behaviors through interactive evolution. In 2016 IEEE Conference on Computational Intelligence and Games (CIG) . IEEE, 1–7

  163. [173]

    Sperrle, M

    F. Sperrle, M. El-Assady, G. Guo, R. Borgo, D. Horng Chau, A. Endert, and D. Keim. 2021. A Survey of Human- Centered Evaluations in Human-Centered Machine Learning. Computer Graphics Forum 40, 3 (2021), 543–568. https://doi.org/10.1111/cgf.14329 arXiv:https://onlinelibrary.wil...

  164. [174]

    Thilo Spinner, Rebecca Kehlbeck, Rita Sevastjanova, Tobias Stähle, Daniel A Keim, Oliver Deussen, and Mennatallah El-Assady. 2024. -generAItor: Tree-in-the-loop Text Generation for Language Model Explainability and Adaptation. ACM Transactions on Interactive Intelligent System...

  165. [175]

    Kaushik Subramanian, Charles L Isbell Jr, and Andrea L Thomaz. 2016. Exploration from demonstration for interactive reinforcement learning. In Proceedings of the 2016 international conference on autonomous agents & multiagent systems . 447–456

  166. [176]

    Prasann Singhal, Tanya Goyal, Jiacheng Xu, and Greg Durrett. 2023. A long way to go: Investigating length correlations in rlhf. arXiv preprint arXiv:2310.03716 (2023)

  167. [177]

    Sumers, Robert D

    Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths, and Dylan Hadfield-Menell. 2022. How to talk so AI will learn: Instructions, descriptions, and autonomy. http://arxiv.org/abs/2206.07870 arXiv:2206.07870 [cs]

  168. [178]

    Sumers, Robert D

    Theodore R. Sumers, Robert D. Hawkins, Mark K. Ho, Thomas L. Griffiths, and Dylan Hadfield-Menell. 2022. Linguistic communication as (inverse) reward design. (April 2022). http://arxiv.org/abs/2204.05091 arXiv: 2204.05091

  169. [179]

    Theodore R Sumers, Mark K Ho, Robert D Hawkins, Karthik Narasimhan, and Thomas L Griffiths. 2021. Learning rewards from linguistic feedback. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 6002–6010

  170. [180]

    Richard S Sutton and Andrew G Barto. 2018. Reinforcement Learning: An Introduction. A Bradford Book, Cambridge, MA, USA

  171. [181]

    Niket Tandon, Aman Madaan, Peter Clark, and Yiming Yang. 2022. Learning to repair: Repairing model output errors after deployment using a dynamic memory of feedback. In Findings of the Association for Computational Linguistics: NAACL 2022, Marine Carpuat, Marie-Catherine de Ma...

  172. [182]

    Matthew E Taylor, Halit Bener Suay, and Sonia Chernova. 2011. Integrating reinforcement learning with human demonstrations of varying ability. In The 10th International Conference on Autonomous Agents and Multiagent Systems- Volume 2. 617–624

  173. [183]

    Bellemare, Jeff Clune, and Joel Lehman

    Felipe Petroski Such, Vashisht Madhavan, Rosanne Liu, Rui Wang, Pablo Samuel Castro, Yulun Li, Jiale Zhi, Ludwig Schubert, Marc G. Bellemare, Jeff Clune, and Joel Lehman. 2019. An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents. arX...

  174. [184]

    Do this instead

    Christopher Thierauf, Ravenna Thielstrom, Bradley Oosterveld, Will Becker, and Matthias Scheutz. 2023. “Do this instead”–Robots that Adequately Respond to Corrected Instructions. ACM Transactions on Human-Robot Interaction (2023)

  175. [185]

    Andrea L Thomaz and Cynthia Breazeal. 2007. Asymmetric interpretations of positive and negative human feedback for a social learning agent. In RO-MAN 2007-The 16th IEEE International Symposium on Robot and Human Interactive Communication. IEEE, 720–725

  176. [186]

    Thomaz and Cynthia Breazeal

    Andrea L. Thomaz and Cynthia Breazeal. 2008. Teachable robots: Understanding human teaching behavior to build more effective robot learners. Artificial Intelligence 172, 6 (2008), 716–737. https://doi.org/10.1016/j.artint.2007.09.009

  177. [187]

    Andrea Lockerd Thomaz, Cynthia Breazeal, et al. 2006. Reinforcement learning with human teachers: Evidence of feedback and guidance with implications for learning performance. In Aaai, Vol. 6. Boston, MA, 1000–1005

  178. [188]

    Andrea Lockerd Thomaz, Guy Hoffman, and Cynthia Breazeal. 2005. Real-time interactive reinforcement learning for robots. In AAAI 2005 workshop on human comprehensible machine learning , Vol. 3. Citeseer

  179. [189]

    Faraz Torabi, Garrett Warnell, and Peter Stone. 2018. Behavioral cloning from observation. arXiv preprint arXiv:1805.01954 (2018). , Vol. 1, No. 1, Article . Publication date: February 2025. Mapping out the Space of Human Feedback for Reinforcement Learning 49

  180. [190]

    Ana C Tenorio-Gonzalez, Eduardo F Morales, and Luis Villasenor-Pineda. 2010. Dynamic reward shaping: training a robot by voice. In Advances in Artificial Intelligence–IBERAMIA 2010: 12th Ibero-American Conference on AI, Bahía Blanca, Argentina, November 1-5, 2010. Proceedings ...

  181. [191]

    Susanne Trick, Franziska Herbert, Constantin A Rothkopf, and Dorothea Koert. 2022. Interactive reinforcement learning with Bayesian fusion of multimodal advice. IEEE Robotics and Automation Letters 7, 3 (2022), 7558–7565

  182. [192]

    Sanne Van Waveren, Christian Pek, Jana Tumova, and Iolanda Leite. 2022. Correct me if I’m wrong: Using non-experts to repair reinforcement learning policies. In 2022 17th ACM/IEEE International Conference on Human-Robot Interaction (HRI). IEEE, 493–501

  183. [193]

    Vivek Veeriah, Patrick M Pilarski, and Richard S Sutton. 2016. Face valuing: Training user interfaces with facial expressions and reinforcement learning. arXiv preprint arXiv:1606.02807 (2016)

  184. [194]

    Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri. 2019. Programmati- cally Interpretable Reinforcement Learning. arXiv:1804.02477 [cs.LG]

  185. [195]

    Chunyang Wang, Yanmin Zhu, Haobing Liu, Tianzi Zang, Jiadi Yu, and Feilong Tang. 2022. Deep Meta-learning in Recommendation Systems: A Survey. arXiv preprint arXiv:2206.04415 (2022)

  186. [196]

    Guanzhi Wang, Yuqi Xie, Yunfan Jiang, Ajay Mandlekar, Chaowei Xiao, Yuke Zhu, Linxi Fan, and Anima Anandkumar

  187. [197]

    Marcel Torne, Max Balsells, Zihan Wang, Samedh Desai, Tao Chen, Pulkit Agrawal, and Abhishek Gupta. 2023. Bread- crumbs to the Goal: Goal-Conditioned Exploration from Human-in-the-Loop Feedback.arXiv preprint arXiv:2307.11049 (2023)

  188. [198]

    Xiaofei Wang, Kimin Lee, Kourosh Hakhamaneshi, Pieter Abbeel, and Michael Laskin. 2022. Skill preferences: Learning to extract and execute robotic skills from human feedback. In Conference on Robot Learning . PMLR, 1259–1268

  189. [199]

    Zhaodong Wang and Matthew E. Taylor. 2019. Interactive Reinforcement Learning with Dynamic Reuse of Prior Knowledge from Human and Agent Demonstrations. InProceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19. International Joint ...

  190. [200]

    Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone. 2017. Deep TAMER: Interactive Agent Shaping in High-Dimensional State Spaces. 32nd AAAI Conference on Artificial Intelligence, AAAI 2018 (Sept. 2017), 1545–1553. https://doi.org/10.48550/arxiv.1709.10163 arX...

  191. [201]

    Garrett Warnell, Nicholas Waytowich, Vernon Lawhern, and Peter Stone. 2018. Deep tamer: Interactive agent shaping in high-dimensional state spaces. In Proceedings of the AAAI conference on artificial intelligence , Vol. 32

  192. [202]

    Nils Wilde, Erdem Bıyık, Dorsa Sadigh, and Stephen L Smith. 2021. Learning reward functions from scale feedback. arXiv preprint arXiv:2110.00284 (2021)

  193. [203]

    Nils Wilde, Alexandru Blidaru, Stephen L Smith, and Dana Kulić. 2020. Improving user specifications for robot behavior through active preference learning: Framework and evaluation.The International Journal of Robotics Research 39, 6 (2020), 651–667

  194. [204]

    Nialah Jenae Wilson-Small, David Goedicke, Kirstin Petersen, and Shiri Azenkot. 2023. A Drone Teacher: Designing Physical Human-Drone Interactions for Movement Instruction. In Proceedings of the 2023 ACM/IEEE International Conference on Human-Robot Interaction . 311–320

  195. [205]

    Junpeng Wang, Liang Gou, Han-Wei Shen, and Hao Yang. 2018. DQNViz: A Visual Analytics Approach to Understand Deep Q-Networks. IEEE Trans Vis Comput Graph (Sept. 2018)

  196. [206]

    Zeqiu Wu, Yushi Hu, Weijia Shi, Nouha Dziri, Alane Suhr, Prithviraj Ammanabrolu, Noah A Smith, Mari Ostendorf, and Hannaneh Hajishirzi. 2023. Fine-Grained Human Feedback Gives Better Rewards for Language Model Training. arXiv preprint arXiv:2306.01693 (2023)

  197. [207]

    Duo Xu, Mohit Agarwal, Ekansh Gupta, Faramarz Fekri, and Raghupathy Sivakumar. 2021. Accelerating Reinforcement Learning using EEG-based implicit human feedback. Neurocomputing 460 (2021), 139–153

  198. [208]

    Kelvin Xu, Zheyuan Hu, Ria Doshi, Aaron Rovinsky, Vikash Kumar, Abhishek Gupta, and Sergey Levine. 2023. Dexterous manipulation from images: Autonomous real-world rl via substep guidance. In 2023 IEEE International Conference on Robotics and Automation (ICRA) . IEEE, 5938–5945

  199. [209]

    Eric Yeh, Melinda Gervasio, Daniel Sanchez, Matthew Crossley, and Karen Myers. 2018. Bridging the gap: Converting human advice into imagined examples. Advances in Cognitive Systems 6 (2018), 1168–1176

  200. [210]

    Lance Ying, Tan Zhi-Xuan, Vikash Mansinghka, and Joshua B Tenenbaum. 2023. Inferring the goals of communicating agents from actions and instructions. In Proceedings of the AAAI Symposium Series , Vol. 2. 26–33

  201. [211]

    J Yow, Neha Priyadarshini Garg, Manoj Ramanathan, Wei Tech Ang, et al. 2024. ExTraCT–Explainable Trajectory Corrections from language inputs using Textual description of features. arXiv preprint arXiv:2401.03701 (2024)

  202. [212]

    Thumbs Up

    Hang Yu, Reuben M Aronson, Katherine H Allen, and Elaine Schaertl Short. 2023. From “Thumbs Up” to “10 out of 10”: Reconsidering Scalar Feedback in Interactive Reinforcement Learning. In 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 41...

  203. [213]

    Christian Wirth, Riad Akrour, Gerhard Neumann, Johannes Fürnkranz, et al. 2017. A survey of preference-based reinforcement learning methods. Journal of Machine Learning Research 18, 136 (2017), 1–46

  204. [214]

    Wenhao Zhan, Masatoshi Uehara, Wen Sun, and Jason D Lee. 2023. How to Query Human Feedback Efficiently in RL? Interactive Learning with Implicit Human Feedback Workshop at ICML 2023 (2023)

  205. [215]

    David Zhang, Micah Carroll, Andreea Bobu, and Anca Dragan. 2022. Time-Efficient Reward Learning via Visually Assisted Cluster Ranking. (Nov. 2022). http://arxiv.org/abs/2212.00169 arXiv:2212.00169 [cs]

  206. [216]

    Jenny Zhang, Samson Yu, Jiafei Duan, and Cheston Tan. 2023. Good Time to Ask: A Learning Framework for Asking for Help in Embodied Visual Navigation. In 2023 20th International Conference on Ubiquitous Robots (UR) . IEEE, 503–509

  207. [217]

    Qiping Zhang, Austin Narcomey, Kate Candon, and Marynel Vázquez. 2023. Self-Annotation Methods for Aligning Implicit and Explicit Human Feedback in Human-Robot Interaction. In Proceedings of the 2023 ACM/IEEE International Conference on Human-Robot Interaction . 398–407

  208. [218]

    Ruohan Zhang, Dhruva Bansal, Yilun Hao, Ayano Hiranaka, Jialu Gao, Chen Wang, Roberto Martín-Martín, Li Fei-Fei, and Jiajun Wu. 2023. A Dual Representation Framework for Robot Learning with Human Guidance. In Conference on Robot Learning. PMLR, 738–750

  209. [219]

    Ruohan Zhang, Faraz Torabi, Lin Guan, Dana H Ballard, and Peter Stone. 2019. Leveraging human guidance for deep reinforcement learning tasks. arXiv preprint arXiv:1909.09906 (2019)

  210. [220]

    Ruohan Zhang, Calen Walshe, Zhuode Liu, Lin Guan, Karl Muller, Jake Whritner, Luxin Zhang, Mary Hayhoe, and Dana Ballard. 2020. Atari-head: Atari human eye-tracking and demonstration dataset. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 6811–6820

  211. [221]

    Wenhao Yu, Nimrod Gileadi, Chuyuan Fu, Sean Kirmani, Kuang-Huei Lee, Montse Gonzalez Arenas, Hao-Tien Lewis Chiang, Tom Erez, Leonard Hasenclever, Jan Humplik, et al. 2023. Language to Rewards for Robotic Skill Synthesis. arXiv preprint arXiv:2306.08647 (2023)

  212. [222]

    Banghua Zhu, Michael Jordan, and Jiantao Jiao. 2023. Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise Comparisons. In Proceedings of the 40th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 202), Andreas...

  213. [223]

    Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B

    Daniel M. Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B. Brown, Alec Radford, Dario Amodei, Paul F. Christiano, and Geoffrey Irving. 2019. Fine-Tuning Language Models from Human Preferences. CoRR abs/1909.08593 (2019). arXiv:1909.08593 http://arxiv.org/abs/1909.08593

  214. [224]

    block-based

    Matt Zucker, J Andrew Bagnell, Christopher G Atkeson, and James Kuffner. 2010. An optimization approach to rough terrain locomotion. In 2010 IEEE International Conference on Robotics and Automation . IEEE, 3589–3595. , Vol. 1, No. 1, Article . Publication date: February 2025. ...

  215. [229]

    Michelle D Zhao, Reid Simmons, and Henny Admoni. 2023. Learning Human Contribution Preferences in Collaborative Human-Robot Tasks. In Conference on Robot Learning . PMLR, 3597–3618

  216. [1209]

    https://doi.org/10.1145/3357236.3395525 arXiv: 2105.12949 Publisher: Association for Computing Machinery, Inc ISBN: 9781450369749

  217. [1498]

    https://doi.org/10.1016/j.promfg.2020.01.138 29th International Conference on Flexible Automation and Intelligent Manufacturing ( FAIM 2019), June 24-28, 2019, Limerick, Ireland, Beyond Industry 4.0: Industrial Advances, Engineering Education and Intelligent Manufacturing

  218. [2014]

    InThe 23rd IEEE international symposium on robot and human interactive communication

    Learning something from nothing: Leveraging implicit human feedback strategies. InThe 23rd IEEE international symposium on robot and human interactive communication . IEEE, 607–612

  219. [2016]

    Autonomous agents and multi-agent systems 30 (2016), 30–59

    Learning behaviors via human-delivered discrete feedback: modeling implicit feedback strategies to speed up learning. Autonomous agents and multi-agent systems 30 (2016), 30–59

  220. [2021]

    CoRR abs/2112.00861 (2021)

    A General Language Assistant as a Laboratory for Alignment. CoRR abs/2112.00861 (2021). arXiv:2112.00861 https://arxiv.org/abs/2112.00861

  221. [2022]

    arXiv:2212.08073 [cs.CL] https://arxiv.org/abs/2212.08073

    Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073 [cs.CL] https://arxiv.org/abs/2212.08073

  222. [2023]

    arXiv preprint arXiv:2305.16291 (2023)

    Voyager: An open-ended embodied agent with large language models. arXiv preprint arXiv:2305.16291 (2023)

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.