{"id":"ba2b4369-f36b-4d8b-aca0-0c239c50db37","arxiv_id":"2506.06324","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A scoping review of 77 papers on human-agent co-learning and co-adaptation finds that most work claims two-way adaptation, with reinforcement learning and decision-making or trust frameworks dominating.","lead":"This paper reviews 77 studies on how people and AI systems learn and adapt together. It is a useful map of a young field, but its counts are inconsistent and its coding methods are not fully documented.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 84.14% two-way finding is internally inconsistent: the PRISMA flow and Table 1 report 77 included studies, but the Results compute percentages over 82, mixing the final set with 5 one-way papers removed as exclusions.","rationale":"The reader's CONDITIONAL verdict seems right, and my concern does not move it. The broad qualitative pattern — two-way adaptation dominates, reinforcement learning is common, studies are small — is plausible and useful, and the paper candidly notes limitations such as the January 2024 cutoff and small samples. But the review claims PRISMA-ScR adherence and provides quantitative synthesis, and the 77/82 inconsistency sits exactly at the point where RQ1 is answered. Because the five one-way papers appear both as a PRISMA exclusion category and as a counted category in the Results, the reported 84.14% cannot be assigned to a well-defined population. This is directly checkable from the authors' own spreadsheet and is a reporting-level fix, which supports CONDITIONAL rather than REJECT. I partially agree with the reader's weakest assumption: undocumented coding is a genuine threat, but the denominator problem is even more load-bearing because it does not depend on judging whether the coders happened to be reliable. If reconstruction confirms the denominator error, the paper retains value as a qualitative map, but its headline percentages should not be relied on until the counts are reconciled.","tokens_in":23878,"tokens_out":4159,"duration_ms":49518,"concrete_test":"Reconstruct the study set from the raw screening log. Enumerate every paper that reached full-text review (N=92), mark each as included or excluded and record its exclusion reason, and note its adaptation-style code from the extraction spreadsheet. Then: (1) compute the two-way prevalence over the final included set (claimed N=77) and over the 82-paper set used in Results; (2) check whether the 5 'one-sided' papers appear both in the exclusion list and in the Results denominator; (3) re-tally Table 2's reinforcement-learning count, since 28.57% equals 22/77, not the stated 21/77. If 69/82 arises because excluded one-way papers were counted, replace 84.14% with 69/77 = 89.61% or revise the inclusion criteria; if 69/77 is correct, recompute every percentage in Results and Discussion with a consistent N.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a quantitative map: 'among 82 papers reviewed, 69 (84.14%) focused on the two-way adaptation style.' For this claim to be sound, the denominator must be the final included set, and the adaptation-style codes must be reproducible. The first condition fails on the face of the manuscript. The PRISMA flow and Table 1 report 77 included studies; the Results section uses 82 as the denominator for adaptation style. The discrepancy is not cosmetic: 69/77 = 89.61%, yet the text states both '69/77 (89.61%)' and '69 (84.14%)' in different places. Moreover, the PRISMA flow lists 'Removed papers with one-sided adaptation style (N=5)' as an exclusion, and the Results' 82-paper set includes 5 one-way papers, i.e., the 84.14% figure counts papers that were supposedly excluded for violating the two-way inclusion criterion. If the final corpus is 77, the headline prevalence is wrong; if the corpus is 82, the inclusion criteria and PRISMA flow are wrong. Similar denominator problems affect the reinforcement learning share: the text says '21/77, 28.57%', but 21/77 is 27.27% and 28.57% equals 22/77. The reader's coding-reliability concern remains real: no codebook, no inter-rater statistic, no released extraction sheets. But the denominator inconsistency is more fundamental, because even a perfect coding scheme cannot make a percentage meaningful when the numerator and denominator come from different study sets.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a PRISMA-ScR scoping review of research on human-agent co-learning and co-adaptation. It states three research questions: RQ1 on terminology used for the relationship, RQ2 on agent types and task domains, and RQ3 on cognitive theories and frameworks. The authors searched Web of Science, Engineering Village, and EBSCOhost, screened 373 records, and report 77 (or sometimes 82) included papers. They present Table 1 characterizing each study and Table 2 listing AI methods and performance metrics; the Results give percentages for adaptation style, cognitive themes, methods, and geographic spread.","tokens_in":24138,"tokens_out":5918,"duration_ms":63465,"significance":"The paper addresses a genuinely emerging topic and would be a useful map if the quantitative claims were reliable. Its contributions include a structured PRISMA flow, two detailed extraction tables, and a first-pass synthesis of terminology. It does not claim a new formal model or derive predictions; its value is empirical synthesis. The main threat is that the headline percentages are internally inconsistent and the coding procedure is not auditable. Because these issues are correctable, the review's underlying contribution remains salvageable.","major_comments":[{"comment":"The central quantitative claim is internally inconsistent. The text states 'among 82 papers reviewed, 69 (84.14%) focused on the two-way adaptation style,' but Figure 2 records 77 included studies and Table 1 is described as covering 77 reports. A later sentence in the same section says '69/77 (89.61%).' The 82-paper denominator appears to include the five one-way papers that Figure 2 lists as removed ('Removed papers with one-sided adaptation style (N=5)'). This is not cosmetic: 69/77 = 89.61%, not 84.14%, and if the inclusion criterion is two-way adaptation, one-way papers cannot be part of the reviewed set. The authors must reconcile the flow, the inclusion criteria, and all percentages using a single consistent corpus.","section":"Results, 'Human-Agent/AI Teams'; Figure 2; Table 1"},{"comment":"The PRISMA flow arithmetic does not add up. From 92 papers for full-text review, the flow says 1 inaccessible and 9 removed by eligibility, i.e., 10 removals, yet reports 77 included (92−10 = 82). If 5 one-sided papers are also removed, removals are 15 and included is 77. The text also says 'the authors removed 10.86% (10/92) of papers,' which is consistent with 82, not 77. These numbers must be corrected and aligned with the final inclusion set.","section":"Results, 'Study Selection' and PRISMA flow"},{"comment":"The most-used-algorithm claim contains an arithmetic error: 'Reinforcement Learning (21/77, 28.57%)' is impossible because 21/77 = 27.27%; 28.57% equals 22/77. The same section's country-region breakdown (Europe 39/77, Asia 26/77, North America 21/77, South America 1/77) sums to 87/77 = 113%, so the geographic percentages are overcounted; the authors need to decide whether multi-country papers are counted once or in each region and recompute.","section":"Results, 'Types of Agents and AI Algorithms' and Discussion"},{"comment":"The coding scheme underlying all prevalence figures is not documented. The manuscript reports Excel notes and consensus discussion but provides no codebook, no operational definition of 'two-way' versus 'one-way' coding (beyond the inclusion statement), no inter-rater reliability statistic, and no release of the extraction sheets. Because every percentage in Results depends on these manual judgments, the classification must be made reproducible (e.g., a codebook, dual-coding counts, and IRR) or the claims should be presented as illustrative rather than quantitative.","section":"Methods, 'Study data collection and synthesis'"}],"minor_comments":[{"comment":"Exclusion criterion #5 is self-contradictory: it is listed under 'We excluded studies with the following criteria' but says 'For comparison, the study included those that state co-adaptation... but describe only the agent's adaptation.' Clarify whether one-way papers were included for comparison, and if so, remove them from the excluded list and align the PRISMA flow.","section":"Methods, 'Inclusion and Exclusion Criteria'"},{"comment":"The citation for the 'first-ever study on mutual adaptation' is given as [66], which in the reference list is Kita et al. on autonomous assistive devices; the text later credits Yamada and Yamaguchi [20] with the 2002 foundation. Correct the reference or reconcile the attribution.","section":"Results, 'Human-Agent/AI Teams'"},{"comment":"The sentence 'Yong describes the mutual adoption of gesture-based human-robot interfaces based on the Wizard of Oz experiment, which is a one-sided approach [19]' appears to mix authors and references; the cited [19] is De Santis (body-machine interfaces), not Yong, and the intended [62] is Xu et al. Check all bracketed citations against the reference list.","section":"Results, 'Concept and functions of Human-Agent/AI Teams'"},{"comment":"Table 1 contains rows with 'NA' entries for adaptation style (e.g., Burke [9] and Damiano/Dumouchel [27]) even though the Results' 69/82 classification implies every row was coded; include a 'not reported' category explicitly in Table 1 and state how NA rows were handled.","section":"Table 1"},{"comment":"The title's 'Human-Agent' and the abstract's 'human-AI-robot' are not clearly distinguished; define the scope (humans with AI and/or robots) in the introduction to avoid terminological drift.","section":"Introduction and title"},{"comment":"There are several typographical and nomenclature errors, including 'Parlo [79]' (should likely be 'Palro'), 'ABIO [20]' (should be 'AIBO'), and 'mutual adoption' (should likely be 'mutual adaptation'). These should be corrected in a careful pass.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"I recommend inviting a revision that requires the authors to recompute all descriptive statistics from a single inclusion list and to provide either a codebook and reliability statistics or a clear downgrading of the quantitative claims. The journal may also request the extraction spreadsheet as a supplementary file. There is no reason to suspect the authors of bad faith; the inconsistencies look like late-stage changes to the inclusion set that were not propagated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a useful literature map with a sloppy numeric foundation. The review aggregates 77 papers on human-agent co-learning/co-adaptation, tabulates agent types, algorithms, cognitive frameworks, and sample sizes, and it lands on a plausible big picture: two-way adaptation dominates, reinforcement learning is the most common method, sample sizes are often tiny, and the field is growing fast. For a newcomer, that map is genuinely helpful.\n\nWhat's new: the aggregation itself. I don't see a framework or measurement contribution, and the \"first-of-its-kind\" claim is asserted rather than demonstrated against prior reviews. But as a scoping review, that's not necessarily a deal-breaker.\n\nThe soft spots are in the numbers. The PRISMA flow and Table 1 say 77 studies were included, but the Results repeatedly use 82 as the denominator—\"69 (84.14%)\" for two-way adaptation. Elsewhere the same paper says \"69/77 (89.61%).\" Both can't be right. The confusion isn't cosmetic: the PRISMA flow lists \"removed papers with one-sided adaptation style (N=5)\" as an exclusion, yet the 82-paper set includes 5 one-way papers. So either the final corpus is 77 and the headline prevalence is wrong, or the corpus is 82 and the inclusion criteria/flow are wrong. There are also arithmetic slips: \"21/77 (28.57%)\" is actually 27.27%. The coding that produces these percentages is also undocumented—no codebook, no inter-rater reliability, no release of the extraction sheets. I can't verify that \"two-way\" means the same thing across papers.\n\nThe \"first-ever\" historical claims (first mutual adaptation study in 2002, first co-learning in 2015) are similarly asserted without checking against earlier reviews. Those should be softened or verified.\n\nThe qualitative map still looks reasonable to me, and the authors do acknowledge small samples and the need for longitudinal work. But the quantitative claims should not be cited as they stand.\n\nBottom line: worth a serious referee, but the authors need to fix the denominators, release the coding materials, and temper the novelty claims. If you're in HRI or human-AI teaming, grab the reference list, but skip the statistics.","headline":"Useful map, shaky numbers: the review's core percentages don't survive contact with its own tables.","tokens_in":24699,"tokens_out":3052,"would_cite":false,"duration_ms":33933,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scoping review of human-agent co-learning finds that two-way adaptation, not one-sided adjustment, dominates the literature.","keywords":["human-agent teaming","co-learning","co-adaptation","mutual adaptation","scoping review","reinforcement learning","human-robot collaboration","cognitive frameworks"],"falsifier":"Re-code a random sample of the 77 included papers with a second coder using the same spreadsheet dimensions, then compute inter-rater agreement; if coders disagree on adaptation style, or if the 84.14% two-way share moves by more than a few points under re-coding, the headline percentages are not stable enough to build on.","tokens_in":23635,"feed_emoji":"🤖","tokens_out":10051,"duration_ms":94352,"temperature":0.7,"pith_summary":"This paper is a scoping review that sets out to organize a young, fragmented research area: humans and intelligent agents that learn from and adapt to each other. The authors screened records from three scholarly databases, included 77 studies, and coded each for adaptation style, agent type, task domain, cognitive framework, and performance measure. Their central finding is that two-way adaptation, where both human and agent adjust, dominates the literature, appearing in 84.14% of the coded papers, with reinforcement learning the most common agent method. They also claim that the vocabulary of 'co-learning,' 'co-adaptation,' 'mutual learning,' and related terms is used inconsistently, so the review's contribution is partly a shared map and partly a call for clearer definitions.","feed_headline":"84% of human-agent studies describe two-way adaptation","feed_subtitle":"A review of 77 studies finds two-way adaptation dominates and reinforcement learning leads agent methods.","key_machinery":"The analytical engine is the adaptation-style classification: each study is assigned to two-way, one-way, both, or unspecified adaptation, and the results are counts of papers under those codes. Around that core code the authors also classify each study by intelligent-agent method, task domain, cognitive framework, and reported performance metric, using a spreadsheet-based synthesis with consensus discussion to resolve disagreements. The two-way code does the load-bearing work because it operationalizes the review's inclusion criterion: a paper counts as co-learning or co-adaptation only when both partners, not just the agent, are adjusting.","core_discovery":"On its own terms, the paper claims to be the first scoping review of human-agent co-learning and co-adaptation, and it claims that the literature is real but terminologically unsettled. It finds that 69 of 82 coded papers (84.14%) describe two-way adaptation, 5 describe one-way adaptation, 4 describe both, and 4 do not specify; that reinforcement learning is the most common intelligent-agent method, reported at 28.57%; and that decision-making, performance, trust, and mental models are the dominant cognitive themes. The paper also traces the field's timeline, crediting a 2002 mutual mind-reading study as the starting point and locating the first use of 'co-learning' in 2015, with the main growth spurt after 2021. If these claims are right, the field's center of gravity is mutual, two-way adjustment, and future systems and measurements should be built around that dyadic loop.","pith_inferences":["The coding may compress a continuum into categories: a paper labeled 'two-way' could still show asymmetric rates, with the human adapting far more than the agent, so future work could measure directionality rather than mere presence.","Because most included studies are lab-based with small samples, the map describes the literature, not the field's real-world effectiveness; transferring these patterns into deployed systems remains untested.","A testable extension: rerun the same search and coding on publications from 2024 onward to see whether the 84.14% two-way share and the reinforcement-learning dominance hold as generative AI enters the space."],"forward_implications":["If the 84.14% two-way finding holds, co-learning and co-adaptation are the modal object of study, so future frameworks should target mutual, not one-sided, adjustment.","If reinforcement learning is truly the most common agent method, progress in RL-based co-learning tools will disproportionately shape the field's next phase.","If decision-making, trust, mental models, and performance are the dominant cognitive themes, evaluation instruments for co-learning should measure all four to be comparable with prior work.","If the terminology is as inconsistent as the review suggests, a shared ontology of 'co-' terms would unify reporting and enable future meta-analyses."],"supporting_citations":[{"why":"Supplies the definitional anchor: co-learning as iterative co-adaptation plus communication, used to identify two-way interaction patterns.","marker":"[39]"},{"why":"Provides the working definition of mutual adaptation between an agent and a human as two-way adjustment toward informed decisions and trust.","marker":"[2]"},{"why":"Identifies challenges of human-AI co-learning and grounds the review's claim that terminology is inconsistent.","marker":"[38]"},{"why":"Supplies a design-pattern method for human-AI co-learning evaluated in a search-and-rescue task, used as an example of two-way co-learning.","marker":"[81]"},{"why":"Documents tangible co-adaptation behaviors in human-robot teams, used to set the inclusion criterion for two-way adaptation.","marker":"[1]"},{"why":"Shows an adaptive agent that complements human policies to optimize team performance, counted as a two-way adaptation study.","marker":"[40]"},{"why":"Presents the bounded-memory formalism for human-robot mutual adaptation, a recurring model across the included studies.","marker":"[14]"},{"why":"Provides a game-theoretic account of human adaptation in human-robot collaboration, used to show how mutual adaptation is modeled.","marker":"[64]"},{"why":"Credited by the review as the first source to introduce the term 'co-learning' in 2015.","marker":"[69]"}],"fun_headline_variants":["84% of human-agent studies are two-way, first review finds","First scoping review: 84% of human-agent studies show two-way adaptation","Scoping review: 84% of human-agent studies assume mutual adaptation","Reinforcement learning dominates methods in human-agent co-adaptation studies","Co-learning vs co-adaptation: terminology unclear in human-agent studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The quantitative results presuppose that the authors' manual assignment of adaptation style and cognitive focus to each paper is reliable and consistent; the paper reports consensus discussion but no inter-rater reliability statistic and no public codebook.","fun_headline_variants_meta":{"raw":{"variants":["84% of human-agent studies are two-way, first review finds","First scoping review: 84% of human-agent studies show two-way adaptation","Scoping review: 84% of human-agent studies assume mutual adaptation","Reinforcement learning dominates methods in human-agent co-adaptation studies","Co-learning vs co-adaptation: terminology unclear in human-agent studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001614,"raw_usage":{"total_tokens":6456,"prompt_tokens":1010,"completion_tokens":5446,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":626,"completion_tokens_details":{"reasoning_tokens":5349}},"tokens_in":626,"tokens_out":5446,"duration_ms":39449,"temperature":1.0,"reasoning_tokens":5349,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:31:56.041990+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-code a random sample of the 77 included papers with a second coder using the same spreadsheet dimensions, then compute inter-rater agreement; if coders disagree on adaptation style, or if the 84.14% two-way share moves by more than a few points under re-coding, the headline percentages are not stable enough to build on.","supporting_citations":[{"cited_title":"Becoming Team Members: Identifying Interaction Patterns of Mutual Adaptation for Human-Robot Co-Learning,","cited_arxiv_id":null,"evidence_quote":"Supplies the definitional anchor: co-learning as iterative co-adaptation plus communication, used to identify two-way interaction patterns."},{"cited_title":"Human-robot mutual adaptation in collaborative tasks: Models and experiments,","cited_arxiv_id":null,"evidence_quote":"Provides the working definition of mutual adaptation between an agent and a human as two-way adjustment toward informed decisions and trust."},{"cited_title":"Six Challenges for Human-AI Co-learning,","cited_arxiv_id":null,"evidence_quote":"Identifies challenges of human-AI co-learning and grounds the review's claim that terminology is inconsistent."},{"cited_title":"Identifying Interaction Patterns of Tangible Co-Adaptations in Human-Robot Team Behaviors,","cited_arxiv_id":null,"evidence_quote":"Documents tangible co-adaptation behaviors in human-robot teams, used to set the inclusion criterion for two-way adaptation."},{"cited_title":"Formalizing Human-Robot Mutual Adaptation: A Bounded Memory Model","cited_arxiv_id":null,"evidence_quote":"Presents the bounded-memory formalism for human-robot mutual adaptation, a recurring model across the included studies."},{"cited_title":"Köppen, Machine Intelligence Research Labs, Kyūshū Daigaku","cited_arxiv_id":null,"evidence_quote":"Credited by the review as the first source to introduce the term 'co-learning' in 2015."}],"review_version":1}