{"id":"7dc11fc2-7a32-47a3-9d73-8a2ff959592d","arxiv_id":"2411.13410","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of prior work on using human and LLM feedback to improve reinforcement learning, plus attention-based methods for large state spaces.","lead":"This paper reviews how human feedback and large language model feedback can help reinforcement learning agents learn faster in complex environments. It also surveys attention-based methods that handle large observation spaces, giving readers a catalog of recent approaches without adding new experiments.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy's reliability rests on unverified ChatGPT-generated paper summaries and an undocumented paper-selection process; an independent audit of a sample of summaries is needed before the survey can be used as a reference.","rationale":"I read the manuscript as a survey whose only evidence is its literature coverage and categorization; there is no new algorithm or result to verify. The reader's weakest assumption—accuracy of ChatGPT-created summaries and absence of a selection protocol—is the same load-bearing point I find. I also note the internal broken cross-reference and mislabeled figure as corroborating signals of insufficient editorial verification, not as attacks on the authors. I do not see reason to move beyond the reader's CONDITIONAL verdict: the audit could either ground the taxonomy or require a revision, so 'UNCHANGED' is the appropriate recommendation. I would not raise the verdict to ACCEPT without a passing audit, and I would not REJECT because the survey may be salvageable and the core concern is empirically testable.","tokens_in":17787,"tokens_out":5490,"duration_ms":59374,"concrete_test":"Draw a pre-specified random sample of 20 of the 83 references. Two independent annotators, blind to each other, compare each survey summary to the original paper's abstract and methods section and code three binary questions: (1) Is the described mechanism present and accurately stated? (2) Is the taxonomy category assigned by the survey consistent with the paper's actual contribution? (3) Does the claimed result appear in the paper? Pre-register a threshold, e.g., more than 2 material inaccuracies or more than 1 wrong category in 20 would fail the survey. This test directly settles whether the AI-generated summaries are reliable enough to support the taxonomy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a literature taxonomy: human/LLM feedback for RL (§2–3) and attention for large observation spaces (§4). That contribution is only as sound as the per-paper summaries and the completeness of the coverage. The Acknowledgment explicitly states that ChatGPT 'create[d] summaries of the cited works,' and the manuscript contains unedited artifacts: Section 1 points to 'Section ??', Section 3.1.1 cites 'Sections 2.1.4 2.1.2 and 2.1.2', and Figure 6 in §2.2.4 is captioned 'Section 2.1.3'. These are not damaging by themselves, but they show that AI-generated content was not carefully verified. More substantively, the survey gives no inclusion/exclusion criteria, so the §4 'large observation space' cluster is supported by only seven papers, several of which (VLN papers in §4.4) are not primarily framed around the curse of dimensionality. If even a small fraction of the ChatGPT-generated summaries misstate a method, a category assignment, or a claimed result, the resulting taxonomy—and any conclusion about what existing work does—cannot be relied on. This is load-bearing because the survey offers no independent check and no reproducible selection protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is a survey of two bodies of RL literature: (i) augmenting RL agents with human or LLM feedback (§2–3), and (ii) attention mechanisms for RL in large observation spaces (§4). The survey proposes a taxonomy, with Section 2 dividing human-feedback work by natural-language versus other feedback, Section 3 dividing LLM-for-RL work into post-ChatGPT and pre-ChatGPT categories, and Section 4 clustering attention-based work into visual attention, curse of dimensionality, interpretability, and visual-language navigation. It closes with a discussion of research gaps and a conclusion. The authors disclose in the Acknowledgment that ChatGPT was used to create summaries of the cited works.","tokens_in":18021,"tokens_out":11276,"duration_ms":109338,"significance":"If reliable, the survey's taxonomy would provide a useful map of the growing human/LLM-feedback literature and the attention-for-large-observation-space literature. The paper does a service by collecting a broad set of recent preprints, by explicitly separating external LLM feedback from LLMs embedded as components, and by flagging open problems such as multi-granular feedback and agent autonomy. However, the contribution rests on the accuracy of per-paper summaries and on the completeness and traceability of the selection process; the disclosed use of ChatGPT for summaries, the absence of selection criteria, and several demonstrable misclassifications mean the survey cannot currently be used as a trusted reference.","major_comments":[{"comment":"The survey gives no stated inclusion/exclusion criteria, search protocol, or selection rationale. The abstract promises a survey 'dedicated to addressing the intricacies of environments characterized by large observation space,' but §4.2 contains only four papers, several of which (e.g., [76] vehicle routing, [77] sequence modeling) are not primarily framed around the curse of dimensionality, while §4.4 includes RT-2 [82], a vision-language-action manipulation model rather than a VLN navigation agent. In addition, Section 2.2.4 on 'Human Feedback as Reward' omits the canonical preference-based RLHF line (e.g., Christiano et al. 2017, Stiennon et al. 2020, Ouyang et al. 2022) even though the paper cites survey [9] on this topic. Without a documented selection process, the reader cannot distinguish a representative survey from an ad-hoc sample.","section":"Section 4 and overall survey scope"},{"comment":"The Acknowledgment states that ChatGPT 'create[d] summaries of the cited works.' A spot-check of the summaries against the cited papers reveals clear misclassifications that affect the taxonomy. In §2.1.2, the summary of [25] ('Yell at Your Robot') describes an 'RL agent' performing action selection and behavior cloning, but that paper is a language-conditioned imitation/diffusion-policy method and does not use reinforcement learning. In §4.4, [82] (RT-2) is presented as part of the Visual-Language-Navigation cluster, but RT-2 is a vision-language-action model for robotic manipulation trained by supervised learning. These are not cosmetic issues; they change category assignments and therefore the survey's substantive claims about families of methods. The authors should audit every summary, correct or remove inaccurate attributions, and describe the verification procedure.","section":"Acknowledgment; §2.1.2; §4.4"},{"comment":"The subsection label 'Human Feedback as Demonstrations' and the blanket statement 'A term for these approaches is called inverse RL' do not fit all papers placed there. For example, [47] uses a Gaussian-process teacher-advice mechanism built on demonstrations, and [46] relies on human intervention and mentoring rather than demonstration-based inverse RL. The criteria for membership in this category should be stated explicitly, and the category name or the placements should be adjusted so that the taxonomy boundaries are testable.","section":"§2.2.5"}],"minor_comments":[{"comment":"The sentence 'The remainder of this paper is organized as follows: Section ?? provides fundamental concepts surrounding RL and LLMs' contains an unresolved placeholder; no such section exists in the manuscript, so either add the promised background section or delete the sentence.","section":"Section 1"},{"comment":"Figure 6, which is placed in §2.2.4, is captioned 'Abstract idea and architecture of papers in Section 2.1.3'; the caption should refer to Section 2.2.4.","section":"Figure 6"},{"comment":"The bullet list says these papers are 'similar to the papers from Sections 2.1.4 2.1.2 and 2.1.2'; the section numbers are garbled and the duplication suggests it should read '2.1.1, 2.1.2, and 2.1.4.'","section":"Section 3.1.1"},{"comment":"There are numerous typos and grammatical errors, including 'preforms' (§2.1.2), 'behaior' (§2.2.1), 'imrpove' (§2.2.1), 'alongisde' (§2.2.1), 'fucntion' (§2.2.5), 'appraoch' (§2.1.3), 'costy' (§3.1.1), 'autnomously' (§3.1.1), 'pretained' (§3.1.2), and 'granularitites' (§5.1); the manuscript needs a careful proofreading pass.","section":"Throughout"},{"comment":"The sentence 'Since, humans may not always know the task, environment and they might know the optimal behavior and decision-making' is not grammatical; the final clause should be revised to express that humans may not know the optimal behavior.","section":"Section 5.4"},{"comment":"The introduction cites existing surveys [9] and [10] but does not explain what this survey adds beyond them; a short positioning paragraph would clarify the intended contribution.","section":"Section 1"},{"comment":"The summary of [31] says the approach uses a 'contextual banding algorithm'; this appears to be a typo for 'contextual bandit algorithm.'","section":"Section 2.1.4"}],"recommendation":"major_revision","confidential_remarks":"For a journal-format survey, the lack of a documented selection strategy and the disclosed use of ChatGPT to generate summaries make the paper risky as a citable reference in its current form. The taxonomy is a defensible contribution in principle, but it needs an audit of all paper summaries, corrected category assignments, and an explicit methodology section before it can be relied upon. The manuscript also reads as an unreviewed preprint with several unedited artifacts; I recommend that the editor require a full revision and a verification report on the AI-generated content before considering publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a survey, not a new-result paper, and its value is entirely in whether its per-paper summaries can be trusted. Right now, I would not rely on them.\n\nWhat the paper does well: It organizes a messy literature into a clear two-part taxonomy: human feedback (Section 2), LLM feedback (Section 3), with a separate section on attention and large observation spaces (Section 4). The pre/post-ChatGPT split in Section 3 is a reasonable organizing device, and the discussion section lists real open problems (multi-granularity feedback, agent autonomy, dynamic communication). If the summaries are accurate, this could serve as a quick orientation for a practitioner or new grad student.\n\nThe soft spots are real and load-bearing. The Acknowledgment says ChatGPT was used to create summaries of cited works. That is not disqualifying by itself, but there is no verification step, no stated inclusion/exclusion criteria, and the manuscript contains unedited artifacts: 'Section ??' in the intro, a figure caption that says 'Section 2.1.3' while sitting in Section 2.2.4, and 'Sections 2.1.4 2.1.2 and 2.1.2' in Section 3.1.1. These suggest the AI-generated content was not carefully checked. For a survey, the entire contribution is accuracy of representation; a single misread can put a paper in the wrong bucket and mislead a reader.\n\nAlso, Section 4's coverage is thin—seven papers for the 'large observation space' cluster, and some (e.g., VLN papers) are not primarily about the curse of dimensionality. The discussion section is speculative but honestly framed as such.\n\nWho is this for? A reader who wants a first map of the feedback-and-attention landscape. They should treat the summaries as leads to the originals, not as reliable descriptions.\n\nMy recommendation: if I were the editor, I would send this to review rather than desk reject, precisely so a referee can demand a selection protocol and a summary audit. The skeleton is fine; the fillings need checking. As it stands, I would not cite it or hand it to a student without warning.\n\nSerious thinker: yes—the organization shows clear thought, but the execution falls short on verification.","headline":"A usable but under-audited survey of human/LLM feedback for RL; the taxonomy is sensible, but the ChatGPT-generated summaries and missing selection protocol mean it is not yet a dependable reference.","tokens_in":18512,"tokens_out":2551,"would_cite":false,"duration_ms":26840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey groups the work on making RL agents work in complex environments into two research directions—human or LLM feedback, and attention mechanisms for large observation spaces—and identifies six gaps the field has not yet closed.","keywords":["reinforcement learning","human feedback","large language models","curse of dimensionality","attention mechanisms","natural language feedback","survey","sample inefficiency"],"falsifier":"Select a random sample of the cited papers and read their abstracts; if a substantial share of the survey's one-line descriptions misstates the method or contribution stated in the original abstract, then the taxonomy's categories are not a faithful map of the field.","tokens_in":17591,"feed_emoji":"🗺️","tokens_out":4775,"duration_ms":48764,"temperature":0.7,"pith_summary":"The paper is a survey that tries to organize the growing literature on improving reinforcement learning agents in complex environments. Its central claim is that the field's main fixes fall into two complementary buckets: supplementing the agent with feedback or assistance from humans or large language models, and using attention mechanisms to cope with large observation spaces. The survey builds a taxonomy of recent work within each bucket, distinguishing natural-language from other feedback, real-time from static guidance, and external from integrated language models. It then lists six open problems, including the lack of multi-granularity feedback, limited agent autonomy, and the absence of standard datasets. A reader would care because the paper offers a structured map of what has been tried and where the field is thin.","feed_headline":"Survey maps RL's two escape routes from hard environments","feed_subtitle":"Human and LLM guidance attack sample inefficiency; attention attacks the curse of dimensionality.","key_machinery":"The central object is the survey's taxonomy itself: a two-part classification scheme that divides the literature into human/LLM feedback (Sections 2 and 3) and attention for large observation spaces (Section 4). The taxonomy does the work of the paper by sorting papers into nested clusters, defined by feedback modality, timing, granularity, and the role the language model plays. It is the lens through which the authors identify the six gaps in Section 5, and it is the deliverable a reader would reuse to position future work.","core_discovery":"The authors claim that the challenges of reinforcement learning—sample inefficiency, poor generalization, and the curse of dimensionality in large observation spaces—can be attacked along two largely separate research lines. The first line augments the RL agent with external guidance: human feedback in forms ranging from natural language instructions to demonstrations, and, more recently, feedback from large language models used as reward designers, planners, or even complete agents. The second line builds attention mechanisms directly into the agent so it can focus on the relevant parts of a large observation space. The survey organizes dozens of papers into a hierarchical taxonomy with finer clusters such as natural-language instructions in simulated versus robotic environments, real-time feedback loops, abstraction and description, LLM as a component versus LLM as an agent, and attention for visual navigation. Its conclusion is that these two directions are largely complementary and that six specific limitations currently constrain progress.","pith_inferences":["The taxonomy could double as a benchmark selection guide: by clustering papers according to feedback type and environment, it implicitly suggests which baseline systems to compare against when testing a new human-in-the-loop or LLM-assisted RL method.","Because the paper's one-sentence summaries of cited works were produced with an AI-based tool and no systematic selection protocol is reported, the finer claims of each cluster would need independent verification against the original papers before being used to support a new research program.","If the identified gaps are real, the field may converge on multi-granularity, adaptive feedback systems that decide when to ask for help—a direction that is only implicit in the survey's discussion.","The survey's split between human feedback and LLM feedback, while useful, may blur as LLM-based systems increasingly stand in for human annotators; a testable extension would examine whether the two clusters yield the same per-paper categorizations when the LLM feedback is itself generated from human demonstrations."],"forward_implications":["New researchers can use the taxonomy as an entry point: first decide whether an approach channels human or LLM feedback, or attacks the observation space with attention, then locate it within the finer clusters.","The six identified gaps—multi-granularity feedback, agent autonomy, long descriptions and instructions, human awareness of tasks, dynamic communication, and the need for datasets—name concrete places where the field is thin.","Combining external feedback with attention mechanisms is a natural next step, since the survey frames the two as complementary responses to the same core challenges.","The sub-categorization by feedback format (natural language versus other, static versus real-time, extrinsic versus integrated) provides a shared vocabulary for comparing otherwise disparate papers.","The emphasis on sample efficiency suggests that benchmark comparisons should report not only final performance but also training-time savings when feedback or attention is added."],"supporting_citations":[{"why":"Anchors the human-feedback branch of the taxonomy as the general reference for reinforcement learning from human feedback.","marker":"[9]"},{"why":"Anchors the LLM-for-RL branch as a prior RL/LLM taxonomy that the survey extends and reorganizes.","marker":"[10]"},{"why":"Supplies the framing of externally influenced agents, the conceptual bridge between human and LLM assistance.","marker":"[11]"},{"why":"Motivates treating natural language as an abstraction for hierarchical RL, a key premise for several feedback sub-categories.","marker":"[12]"},{"why":"Introduces the curse-of-dimensionality framing used to define the survey's second focus area on large observation spaces.","marker":"[8]"},{"why":"Provides the grounding for the dynamic and real-time feedback clusters as a review of interactive RL from human social feedback.","marker":"[5]"}],"fun_headline_variants":["RL's two escape routes: human/LLM feedback and attention","How feedback and attention help RL tackle large observation spaces","Survey: RL aided by human/LLM feedback and attention","RL's sample inefficiency meets two fixes: feedback and attention","Attention and external feedback: RL's cure for big spaces"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire taxonomy depends on the correctness of the one-sentence summaries of each cited paper, which the authors disclose were produced with an AI-based tool; if any of those summaries misrepresents the original work, the classification built on them is unreliable.","fun_headline_variants_meta":{"raw":{"variants":["RL's two escape routes: human/LLM feedback and attention","How feedback and attention help RL tackle large observation spaces","Survey: RL aided by human/LLM feedback and attention","RL's sample inefficiency meets two fixes: feedback and attention","Attention and external feedback: RL's cure for big spaces"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001333,"raw_usage":{"total_tokens":5416,"prompt_tokens":935,"completion_tokens":4481,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":551,"completion_tokens_details":{"reasoning_tokens":4397}},"tokens_in":551,"tokens_out":4481,"duration_ms":30110,"temperature":1.0,"reasoning_tokens":4397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:25:16.612704+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of the cited papers and read their abstracts; if a substantial share of the survey's one-line descriptions misstates the method or contribution stated in the original abstract, then the taxonomy's categories are not a faithful map of the field.","supporting_citations":[{"cited_title":"A conceptual framework for externally-influenced agents: An assisted reinforcement learning review","cited_arxiv_id":null,"evidence_quote":"Supplies the framing of externally influenced agents, the conceptual bridge between human and LLM assistance."},{"cited_title":"Reinforcement learning: A tutorial survey and recent advances","cited_arxiv_id":null,"evidence_quote":"Introduces the curse-of-dimensionality framing used to define the survey's second focus area on large observation spaces."},{"cited_title":"A review on interactive reinforcement learning from human social feedback","cited_arxiv_id":null,"evidence_quote":"Provides the grounding for the dynamic and real-time feedback clusters as a review of interactive RL from human social feedback."}],"review_version":1}