{"id":"7ccb269e-e369-44b7-aec4-45afd7e73117","arxiv_id":"2412.20573","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A single empirical-progress measure, im(σ,ω)=κ(σ)⋅progress(σ,ω), is proposed to drive a robot's choices of tasks, learning strategies, and tutors across reinforcement and imitation learning for simple and sequential tasks.","lead":"This thesis by Sao Mai Nguyen summarizes a research program on robots that choose what, when, and from whom to learn by measuring their own progress. It argues that one intrinsic-motivation formula, based on empirical competence progress, can unify autonomous exploration, imitation, and tutor selection for simple and sequential tasks.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unifying criterion im(σ,ω)=κ(σ)·progress(σ,ω) rests on a cost coefficient that Section 5.3 admits was set arbitrarily, so the claimed robustness and efficiency are conditional on an unvalidated trade-off parameter.","rationale":"The reader's weakest-assumption analysis identifies the same load-bearing concern: the arbitrary cost coefficient κ(σ) in Eqs. 3.1 and 5.1. This is the right single point because the thesis's central claim is not merely that progress is a useful signal, but that one product form im=κ·progress unifies all strategy and task choices. If κ is a free constant, the framework can be tuned to display or hide the claimed advantages, and the thesis offers no evidence that the conclusions are insensitive to it. This is an internally acknowledged limitation, not a disagreement with consensus: Section 5.3 explicitly says κ was set arbitrarily and delegates a principled formulation to future work. The thesis does draw on prior peer-reviewed publications, which is legitimate supporting evidence; the missing piece is not new experiments but an explicit sensitivity analysis of the central selection rule. A κ-sweep on an existing SGIM-PB setup would settle whether the unified criterion's practical behavior depends materially on the arbitrary constant. Since the reader's CONDITIONAL verdict already reflects this limitation, no change in verdict is warranted.","tokens_in":48112,"tokens_out":5143,"duration_ms":58771,"concrete_test":"Re-run the SGIM-PB hierarchical Yumi task from Duminy et al. (2021) with the cost coefficient κ(σ) swept over at least {0.01, 0.1, 1, 10} for both a helpful and an unhelpful teacher, while fixing all other settings. Report final success on the hardest outcome space Ω5, the number of demonstration requests, and the fraction of time spent in autonomous versus imitation strategies. If final competence and query counts are stable across the sweep and the optimal κ does not shift with tutor quality, the concern is weakened; if performance or query count changes materially, the claimed robustness is conditional on an arbitrary parameter, confirming the Section 5.3 limitation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the unified intrinsic-motivation criterion im(σ,ω)=κ(σ)·progress(σ,ω) (Eqs. 3.1 and 5.1) that lets a single learner choose task, strategy, and teacher. The most fragile ingredient is κ(σ). Section 5.3 states that 'in the previous studies, this was considered a constant set arbitrarily' and only promises future work to derive it from teacher availability, willingness, and other motivational drives. This matters because the learner's selection rule compares im values across strategies: κ is not an inert normalizer but the explicit trade-off weight between autonomous exploration, imitation of policies, imitation of goals, and imitation of task decompositions, and between tutors of different reliability and availability. Different κ values can therefore change the learner's curriculum and the number of demonstrations it requests. The headline empirical claims of being 'more robust to the quality of the tutoring' and learning 'faster with fewer demonstrations' are properties of this choice rule, so they inherit the unvalidated parameter. The thesis reports no sweep over κ, no estimation procedure, and no argument that favorable results survive a wide range of κ. Because the thesis itself identifies this as unresolved, the unified formulation is plausible but its claimed advantages are conditional on a constant the document does not justify.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript is the author's Habilitation à Diriger des Recherches thesis, posted on arXiv, synthesizing roughly a decade of work in developmental cognitive robotics. The central claim is that a single intrinsic-motivation criterion, im(σ,ω)=κ(σ)·progress(σ,ω) (Eqs. 3.1 and 5.1), lets a learner choose its own curriculum by actively selecting the task ω, the learning strategy σ (autonomous exploration versus imitation, low-level actions versus task decomposition), and the tutor from whom to request demonstrations. The thesis reviews the SGIM family of algorithms (SGIM-D, SGIM-ACTS, SGIM-PB), the IM-PB and CHIME architectures for compositional tasks, the STAR hierarchical-RL algorithm, and several applications to socially assistive robotics, and it claims that an active learner is more robust to poor tutors and learns faster with fewer demonstrations. No new experiments are reported; the empirical claims are inherited from the cited papers.","tokens_in":48324,"tokens_out":5194,"duration_ms":53636,"significance":"If the unified criterion works as claimed, it would provide a valuable bridge between reinforcement learning, imitation learning, and hierarchical task decomposition, and it would give a concrete computational model of an active learner that selects teachers and learning strategies. The thesis contains a useful structured survey of active imitation learning, a taxonomy (Table 1.4.1), and a coherent synthesis of the author's previously published algorithms, including STAR's reachability-based spatial abstraction, which carries theoretical suboptimality guarantees. The discussion of rehabilitation and ASD coaching also indicates practical impact. However, the document contains no new experiments, and the formal framework leaves the cost coefficient in the central equation as an arbitrary constant; the claimed empirical advantages are therefore not verified within the manuscript. The significance is conditional on a validation that is currently absent.","major_comments":[{"comment":"Eqs. (3.1) and (5.1) define the central criterion im(σ,ω)=κ(σ)·progress(σ,ω), but §5.3 states that in previous studies κ was 'a constant set arbitrarily'. Because the learner selects strategies by comparing im values, κ is not a harmless normalization: it is the explicit trade-off weight between autonomous exploration, imitation of policies, goals, and task decompositions, and between tutors with different availability and reliability. A poorly chosen κ can change the learner's curriculum and the number of demonstrations requested, so the headline claims of robustness and faster learning are conditional on a parameter the thesis does not justify. A revision should include a sensitivity analysis over κ, a principled estimation procedure, or an explicit removal of the empirical claims from the statement of the contribution.","section":"§5.3, Eqs. (3.1) and (5.1)"},{"comment":"The manuscript presents no new experimental data; every empirical assertion in Chapters 3 and 4 is a summary of previously published work. The text reports qualitative outcomes, such as 'SGIM-PB outperforms SGIM-ACTS' in §3.4 and robustness to poor demonstrations in §3.3, without effect sizes, error bars, or comparisons in this document. Consequently, the reader cannot independently verify the central 'more robust and faster' claim. The document should either reproduce key experiments, include quantitative summaries of the underlying papers, or be explicitly framed as a review with the unifying equation as the sole new contribution.","section":"Chapters 3 and 4"},{"comment":"The core notion progress(σ,ω) is never formally defined in the thesis. Section 1.3.1 defines strategies only as data-collection heuristics, and §3.5 says progress is 'measured through the last episodes' without specifying the estimator, the window size, or how competence is computed for hierarchical goals. Since the same progress measure is used both to define intrinsic motivation and to update the interest map that drives the curriculum, the 'common formulation' is hard to falsify without a precise definition. The thesis should give an explicit definition of progress and state the assumptions under which Eq. (3.1) is a valid reward for strategy selection.","section":"§1.3–§3.5"}],"minor_comments":[{"comment":"The sentence 'this is was considered a constant set arbitrarily' contains a typo ('is was') and should read 'this was considered a constant set arbitrarily'.","section":"§5.3"},{"comment":"The phrase 'may the demonstrations requested to teachers be low-level policies, goals or decomposition into subgoals' is ungrammatical; consider 'whether the demonstrations requested from teachers are low-level policies, goals, or decompositions into subgoals'.","section":"§3.5"},{"comment":"The column headers 'Environ.', 'Imitation', and 'Query' are ambiguous, and several cells mix 'Low-level', 'Policy', and 'Outcome' without explaining the taxonomy; adding a legend would improve readability.","section":"Table 1.4.1"},{"comment":"There are typos in the caption: 'ST AR' should be 'STAR' and 'seperated' should be 'separated'.","section":"Fig. 2.2.2"},{"comment":"The line 'Ensure: partition of outcome spaces R ← F i{Ωi}' is not standard notation and should be defined explicitly, including the condition under which a region is split.","section":"Algo. 2.1.1"}],"recommendation":"major_revision","confidential_remarks":"The document is a habilitation thesis rather than a standard journal article. If the venue accepts self-archived theses as research contributions, the lack of new experiments is less problematic, but the arbitrary cost coefficient remains a load-bearing issue. For a regular cs.AI paper, the absence of new experiments and the unvalidated κ are serious. I would advise the editor to treat the unifying equation as a promising research proposal and require either a sensitivity analysis, a principled derivation of κ, or an explicit framing of the manuscript as a synthesis with the empirical claims clearly attributed to earlier papers."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a habilitation synthesis, not a primary research paper. The central formulation im(σ,ω)=κ(σ)·progress(σ,ω) already appears in the author's earlier SGIM papers (Nguyen and Oudeyer 2012d; Duminy et al. 2021), so the thesis adds framing, a literature structure, and a research agenda rather than new math or new experiments. Read as a position or survey document, it is honest and useful; read as a new result, it is not.\n\nWhat it does well: Table 1.4.1 organizes active imitation learning along what/how/when/who in a way I have not seen laid out as compactly elsewhere. The distinctions between mimicry and emulation, and between demonstrations of policies, goals, and task decompositions, are clear and consistent throughout. The STAR reachability-abstraction material in Section 2.2 gives a real theoretical anchor—a suboptimality bound—though that comes from Zadem et al. 2024, not from this manuscript. The Keraal dataset and rehabilitation application show a concrete downstream use. The citation pattern is transparent: the earlier SGIM papers are cited directly, so the thesis is not laundering prior work.\n\nSoft spots: the main one is κ. It is not an inert normalizer; it is the trade-off weight between autonomous exploration, imitation, and different tutors. Section 5.3 admits that in previous studies it was set arbitrarily, and the thesis reports no sweep over κ and no estimation method. The headline claims of robustness to poor tutoring and faster learning with fewer demonstrations are properties of the strategy-choice rule, so they inherit this unvalidated parameter. If κ is poorly chosen, the learner might over-request demonstrations or under-explore, and the claimed advantages would not follow. The thesis is honest about this, but it means the unified formulation is plausible rather than established.\n\nI would not call the self-referential progress measure a circularity burden; using the learner's own competence history to guide exploration is the standard intrinsic-motivation architecture. But it does mean the guarantees are empirical, not formal, and the thesis gives no argument that the same progress measure is uniquely justified for task, strategy, and tutor selection. Finally, the supporting results are summaries of prior papers, so a reader who wants to verify them must go to the originals.\n\nWho it is for: someone entering developmental robotics or interactive RL who wants a compact map of the field and of the author's contributions. It deserves a serious referee if the venue is a survey, position, or synthesis outlet; it should not be treated as a new-contribution research paper. If it goes to review, the main request should be a κ sensitivity analysis or an explicit scoping of the claims to the chosen constant.","headline":"A transparent habilitation synthesis that restates the SGIM progress-based intrinsic-motivation formula from prior papers and is honest about its own main weakness: the cost coefficient κ is an arbitrary constant.","tokens_in":48886,"tokens_out":2625,"would_cite":false,"duration_ms":30400,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single competence-progress measure, weighted by strategy cost, is claimed to unify a robot's choice of task, tutor, and learning strategy.","keywords":["intrinsic motivation","active imitation learning","social guidance","curriculum learning","hierarchical reinforcement learning","multi-task learning","socially assistive robotics","activities of daily living"],"falsifier":"Run the SGIM-PB learner on a fixed hierarchical task set with two tutors, one helpful and one repetitive, while sweeping the autonomous-exploration cost $\\kappa$ from much smaller to much larger than the imitation cost. If learning speed and final competence vary sharply across the sweep, the unification claim is only as strong as the cost calibration; a truly unified formulation should be insensitive to that arbitrary constant.","tokens_in":47874,"feed_emoji":"🤖","tokens_out":10347,"duration_ms":94769,"temperature":0.7,"pith_summary":"This thesis argues that one quantity can drive a robot's entire learning strategy: the recent improvement in the robot's competence on a task, discounted by the cost of the learning method used. The same formula $im(\\sigma,\\omega)=\\kappa(\\sigma)\\cdot progress(\\sigma,\\omega)$ is claimed to cover exploring on its own, asking a tutor for a demonstration, choosing among several tutors, and deciding whether to learn a task directly or break it into subtasks. If the claim holds, a robot becomes an active consumer of human guidance instead of a passive receiver of datasets: it decides when to ask, whom to ask, and what kind of help to request, and it keeps learning effectively even when tutors are imperfect. The practical payoff would be faster learning of sequential, compositional tasks with fewer demonstrations, and the same progress measure can drive a robot coach that personalises exercise curricula for human pupils.","feed_headline":"Same curiosity formula drives exploration and imitation learning","feed_subtitle":"A learner that scores each option by competence progress builds its own curriculum and needs fewer demonstrations","key_machinery":"The load-bearing machinery is the cost-weighted competence-progress identity $im(\\sigma,\\omega)=\\kappa(\\sigma)\\cdot progress(\\sigma,\\omega)$. Here $progress(\\sigma,\\omega)$ is the empirical change in competence for goal $\\omega$ under strategy $\\sigma$ over recent episodes, with competence defined in the multi-task setting as a common reward typically based on the distance between the desired and reached outcome. The identity carries the argument because the same $progress$ term is used to score autonomous action exploration, autonomous outcome-space exploration, autonomous task-decomposition exploration, mimicry of demonstrated actions, emulation of demonstrated goals, and imitation of a demonstrated task decomposition. The cost coefficient $\\kappa(\\sigma)$ converts teacher availability, willingness, and other interaction costs into a comparable scale, so the learner can choose rationally among strategies that consume very different human or physical resources.","core_discovery":"The central discovery is a unification claim: one intrinsic-motivation measure based on empirical progress is valid for both autonomous exploration and social guidance, whether the demonstration requested is a low-level policy, a goal, or a decomposition into subgoals. Formally, the thesis proposes $im(\\sigma,\\omega)=\\kappa(\\sigma)\\cdot progress(\\sigma,\\omega)$, where $\\sigma$ is a learning strategy, $\\omega$ a task or goal outcome, $progress$ is measured over the last episodes on that strategy and goal, and $\\kappa(\\sigma)$ is the cost of the strategy, representing tutor availability and willingness. This turns previously separate decisions — which task to practise, whether to explore or imitate, whether to ask for an action or a goal, and which tutor to consult — into one selection problem in which each (strategy, goal) option is scored by the same estimated reward. The thesis reports that the resulting Socially Guided Intrinsic Motivation (SGIM) algorithms learn multi-task and hierarchical task sets, transfer knowledge across tasks, switch automatically from simple to complex tasks, and remain effective when tutors give poor or repetitive demonstrations. It also argues that the same progress measure supports emerging symbolic task representations, bridging continuous sensorimotor learning and language-like communication with tutors.","pith_inferences":["Extension: because the formula treats imitation and exploration as strategies with a common currency, one could apply it to a single learner that switches between reinforcement learning and behaviour cloning at the level of neural policies, rather than only the low-level continuous control tasks studied here.","Extension: the progress term is a derivative of competence, so the method's practical success should depend on the time scale used to estimate progress; a natural next step is a statistical treatment of progress estimation.","Extension: if the cost coefficient were learned from human coaching data instead of set arbitrarily, the model could predict when a human learner asks for help, connecting to models of help-seeking behaviour in education.","Extension: placing progress-based selection on top of a reachability-based symbolic abstraction, like the one in the STAR algorithm, would yield a tutor-aware version of goal-conditioned hierarchical reinforcement learning."],"forward_implications":["A robot using the progress criterion will automatically order its curriculum: it practises easy tasks early, moves to hierarchical tasks only after their subtasks are mastered, and switches from imitation early in training to autonomous exploration later.","Choosing between mimicry and emulation, and between asking for a policy, a goal, or a task decomposition, becomes an empirical question the learner answers from its own progress data rather than a designer's choice.","With several tutors, the learner weights each tutor by the competence progress that tutor enables, so it can ignore a poor teacher and concentrate requests on the teacher that is expert for each outcome.","The same reward can be applied in an intelligent tutoring system: instead of a fixed exercise schedule, a robot coach selects exercises that maximise each student's progress, and can discover prerequisite relationships between exercises from score data alone."],"supporting_citations":[{"why":"Supplies the typology of intrinsic motivation and the notion of empirical progress that the thesis adopts.","marker":"Oudeyer and Kaplan [2009]"},{"why":"Formalises curiosity and competence progress as an intrinsic reward, giving the progress measure its theoretical grounding.","marker":"Schmidhuber [2010]"},{"why":"Introduces SGIM-ACTS, the first algorithm in which the learner actively chooses teachers, strategies, and goals by competence progress.","marker":"Nguyen and Oudeyer [2012d]"},{"why":"Establishes SGIM-D and shows that combining autonomous exploration with demonstrations improves generalisation over continuous motor tasks.","marker":"Nguyen and Oudeyer [2014]"},{"why":"Extends the progress criterion to demonstrations of task decomposition and hierarchical tasks, the load-bearing case for the sequential-task claim.","marker":"Duminy et al. [2021]"},{"why":"Defines active imitation learning, the problem setting the thesis generalises to multi-task and hierarchical learning.","marker":"Shon et al. [2007]"},{"why":"Provides a parallel algorithm, CLIC, showing the same learning-progress criterion can select what, how, when, and whom to imitate.","marker":"Fournier et al. [2019]"}],"fun_headline_variants":["One progress score unifies exploration, imitation, and tutor choice","Same progress metric drives autonomous and guided learning","Curiosity metric lets robots pick tasks, tutors, and imitation","One intrinsic motivation rule for exploration, imitation, and tutoring","Progress-based curiosity unifies self-learning and tutor requests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on being able to choose, for each way of learning, a number that says how expensive that way is; the thesis admits in Section 5.3 that this number was picked arbitrarily in earlier experiments, and picking it badly could undo the promised gains in speed and robustness.","fun_headline_variants_meta":{"raw":{"variants":["One progress score unifies exploration, imitation, and tutor choice","Same progress metric drives autonomous and guided learning","Curiosity metric lets robots pick tasks, tutors, and imitation","One intrinsic motivation rule for exploration, imitation, and tutoring","Progress-based curiosity unifies self-learning and tutor requests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000677,"raw_usage":{"total_tokens":3137,"prompt_tokens":1064,"completion_tokens":2073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":680,"completion_tokens_details":{"reasoning_tokens":1994}},"tokens_in":680,"tokens_out":2073,"duration_ms":12614,"temperature":1.0,"reasoning_tokens":1994,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:16:33.139923+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the SGIM-PB learner on a fixed hierarchical task set with two tutors, one helpful and one repetitive, while sweeping the autonomous-exploration cost $\\kappa$ from much smaller to much larger than the imitation cost. If learning speed and final competence vary sharply across the sweep, the unification claim is only as strong as the cost calibration; a truly unified formulation should be insensitive to that arbitrary constant.","supporting_citations":[],"review_version":1}