{"id":"9a59d156-e5fb-403f-9cb2-97bd2f921725","arxiv_id":"2607.23515","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"LLMs generate subtask decompositions and ACL task spaces so sparse-reward RL solves long-horizon manipulation better than dense human rewards on five LIBERO tasks.","lead":"LEACL uses large language models to break long robot tasks into subtasks and to invent the parameter spaces that automatic curriculum learning needs, so agents can learn from sparse success/fail rewards only. It matches or beats hand-tuned dense rewards on five simulated kitchen-style manipulation tasks.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The headline \"sparse beats dense\" rests on a dense-reward baseline authored by the same team, asserted to be near-optimal without evidence, and showing instability (0–30% success, ±37–50% CI) that weakens the comparison it supports.","rationale":"Good-faith reading: LEACL is a systems paper whose core empirical payload is Table II, and the boldest element of the strongest claim is the win over dense rewards, not the win over sparse or w/o-ACL baselines (those are near-zero and uncontroversial). The least secure condition for that element is baseline quality. The paper is transparent about most limitations — a priori predicate grammar, sim-only, single ACL backend (ADR), 5 seeds — and the reader correctly flagged the predicate vocabulary as the weakest methodological assumption. My concern is complementary rather than identical, hence \"partial\" agreement: the reader identified the strongest claim and its dependency accurately but did not press on the LEAGUE baseline's evidentiary status, which is the load-bearing piece for the comparative headline. I do not propose a verdict downgrade: the absolute performance, the human-curriculum parity, and the disclosed limitations support CONDITIONAL with medium correctness risk. The additional condition I'd attach is an independent or automated dense-reward comparison before the \"surpasses human-designed dense rewards\" framing is taken at face value. Secondary, lesser issues: (1) only 5 seeds with CIs wide enough that the Task 4 gap vs human curriculum (60.6 vs 79.7) and the Task 5 ranking could shift; (2) LLM variance in decomposition quality is unquantified — results appear to use a single successful GPT-4o-mini decomposition per task; (3) LIBERO+ parameterized predicates are themselves human engineering, so \"without any hand-crafted effort\" (contributions list) is mildly overstated, though Sec. VII concedes this.","tokens_in":13730,"tokens_out":1691,"duration_ms":43426,"concrete_test":"Replace the author-crafted LEAGUE rewards on Tasks 3 and 5 with (a) rewards from an automated dense-reward generator (e.g., Text2Reward/Eureka-style) and (b) dense rewards written by a roboticist not on the author team, run with ≥10 seeds. If LEACL still matches or beats the best dense-reward variant by a significant margin, the headline comparative claim holds; if any dense variant closes the gap, the paper's contribution should be reframed as \"ACL with LLM-generated task specifications\" rather than \"sparse beats dense.\"","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is comparative: LEACL with sparse rewards outperforms \"human-designed dense rewards\" (LEAGUE) on most of five tasks (Table II: 89.4% vs 29.8% ketchup; 75.9% vs 0% mug-microwave). For this claim to land, the LEAGUE baseline must be a genuinely strong dense-reward policy. The paper asserts (Sec. V-C-d) that its handcrafted shaping rewards \"likely represent an upper bound on the performance achievable by a generated reward policy,\" but offers no evidence for this: the rewards were designed by the authors themselves, are not shown, and no external reference point is provided (no automated reward-generation baseline like Text2Reward or Eureka, despite these being the natural comparators discussed in Sec. II). The numbers themselves raise doubt: LEAGUE scores 0.0% on mug-in-microwave and has enormous variance elsewhere (71.0±37.4 on bowl-on-plate, 29.8±50.6 on ketchup — the CI spans nearly the full range over only 5 seeds), which is more consistent with a brittle or poorly tuned baseline than with a stable upper bound. Meanwhile the paper's own narrative (Sec. VI-b) uses LEAGUE's failures to argue dense rewards are hard to design — a claim that is true in general but is here supported by a single internal attempt. If a moderately better dense reward (from an independent expert or an automated generator) substantially closes the gap on ketchup and mug-microwave, the central framing \"sparse beats dense\" collapses into the weaker \"ACL task parameterization helps,\" which the LEACL w/o ACL baseline already establishes differently. Note this is a comparison-validity concern, not a correctness error: LEACL's absolute results are solid and the predicate-vocabulary caveat (the reader's weakest assumption) is honestly disclosed.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful move here is narrow and real. Instead of asking the LLM for dense rewards or a full curriculum, they use it only to decompose the task into PDDL subtasks and to emit the parameter spaces plus difficulty orderings that off-the-shelf ACL needs. Then ADR runs on pure sparse completion rewards. That is the distinctive piece relative to LEAGUE++, Text2Reward, Eureka, etc.\n\nWhat works: the three-stage pipeline is clear, LIBERO+ (parameterized predicates) is a practical extension they ship, and Table II is readable—five seeds, 1000 eval episodes, coherent ablations. LEACL lands high asymptotic success (e.g. ~89% ketchup, ~76% mug-microwave) and sits close to the fully human curriculum while crushing plain sparse and decomposition-without-ACL. They are also honest that the full predicate vocabulary is given a priori; that is the real scope limit, not a hidden one.\n\nSoft spots, in proportion. Only one ACL backend (ADR) is tried, LLM generation variance is not averaged, and everything is sim kitchen. The bigger comparison issue is LEAGUE: the dense rewards are the authors’ own, never shown, asserted as near-upper-bound without external check (no Text2Reward/Eureka run), and the numbers are unstable (0% on mug-microwave, CIs of ±37–50 on other tasks over five seeds). So the framing “sparse beats human dense” is weaker than the abstract suggests; the safer reading is “good task parameterization + ACL makes sparse work,” which their own w/o-ACL ablation already supports. That does not sink the absolute LEACL numbers or the human-curriculum parity.\n\nThis is for people who actually train multi-step manipulation policies and are tired of reward sculpting. Math is standard PPO/ACL; citations are appropriate; no circularity. I would send it to peer review—solid enough systems work for a robotics or robot-learning venue, with the baseline and multi-ACL gaps as revision items. Worth reading if you touch curricula or LLM+RL; not a must-cite outside that lane.","headline":"Clean systems result: LLM meta-specs let sparse-reward ACL nearly match a human curriculum on five LIBERO tasks, but the 'sparse beats dense' headline rests on a brittle internal baseline.","tokens_in":15163,"tokens_out":540,"would_cite":true,"duration_ms":22621,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Language models can supply the task spaces and difficulty orderings that let automatic curriculum learning solve long-horizon robot manipulation from sparse rewards alone.","keywords":["automatic curriculum learning","large language models","long-horizon manipulation","sparse rewards","task decomposition","PDDL","reinforcement learning","LIBERO"],"falsifier":"Retrain the five LIBERO+ tasks with the identical LLM pipeline but replace the generated task spaces by random or empty parameterizations; if success rates then collapse to the near-zero levels of the sparse-reward and no-ACL baselines, the claim that the LLM-generated specifications are doing the essential work is falsified.","tokens_in":14917,"feed_emoji":"🦾","tokens_out":796,"duration_ms":23657,"temperature":0.7,"pith_summary":"Long-horizon robot manipulation is hard for reinforcement learning because success signals arrive only at the very end. Automatic curriculum learning can help by training on easier variants first, but it normally needs humans to hand-craft the space of task variants and a difficulty measure. This paper shows that a large language model can generate both the subtask breakdown and those missing specifications, after which a standard curriculum algorithm learns each subtask from sparse completion rewards only. On five multi-step LIBERO manipulation tasks the resulting method reaches higher final success rates than a strong baseline that uses carefully engineered dense rewards, and approaches a fully human-designed curriculum. The practical payoff is that complex manipulation skills can be acquired with far less manual reward and curriculum engineering.","feed_headline":"LLMs build curricula so robots learn long tasks from sparse rewards","feed_subtitle":"On five LIBERO manipulation benchmarks the method beats hand-designed dense rewards at final success rate","key_machinery":"LEACL’s three-stage pipeline: LLM task decomposition into PDDL subtasks, LLM meta-task generation of control-parameter spaces and difficulty-ordered instances, and plug-and-play automatic curriculum learning that samples those instances under sparse rewards.","core_discovery":"When an LLM first decomposes a long-horizon manipulation goal into PDDL subtasks and then emits a parameterized task space plus an ordered difficulty sequence for each subtask, ordinary automatic curriculum learning can train a single policy to high asymptotic success using only sparse task-completion rewards, outperforming human-designed dense rewards on most of the evaluated tasks.","pith_inferences":["If the predicate-vocabulary bottleneck can be lifted by letting the LLM propose new predicates, the method would become domain-agnostic rather than library-dependent.","The same LLM-generated task spaces could be reused across different base RL algorithms or even imitation-learning curricula, not only PPO-plus-ADR.","Performance gaps that remain versus the human curriculum likely trace to the quality of the difficulty ordering rather than the decomposition itself, suggesting a cheap human-in-the-loop re-ranking step."],"forward_implications":["Dense reward design can be removed from many multi-step manipulation pipelines without sacrificing final success rate.","Existing automatic curriculum algorithms become usable on new robot tasks once an LLM supplies the missing parameter space and difficulty measure.","The same three-stage pattern can be applied to any domain that already possesses a predicate vocabulary and a sparse completion signal.","Human effort shifts from writing reward functions to curating the predicate grammar and the initial natural-language goal."],"fun_headline_variants":["LLMs decompose tasks so ACL trains robots from sparse rewards only","LEACL lets LLMs supply subtask specs for sparse-reward curricula","LLM curricula enable ACL to beat dense rewards on LIBERO tasks","From PDDL subtasks to difficulty order: LLMs drive sparse ACL","Sparse-reward ACL matches dense designs when LLMs set the curriculum"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The full set of predicates that describe objects, fixtures and goals must already be supplied to the language model; the method does not invent new predicates on its own.","fun_headline_variants_meta":{"raw":{"variants":["LLMs decompose tasks so ACL trains robots from sparse rewards only","LEACL lets LLMs supply subtask specs for sparse-reward curricula","LLM curricula enable ACL to beat dense rewards on LIBERO tasks","From PDDL subtasks to difficulty order: LLMs drive sparse ACL","Sparse-reward ACL matches dense designs when LLMs set the curriculum"]},"model":"grok-4.5","effort":"low","cost_usd":0.004029,"raw_usage":{"total_tokens":1286,"prompt_tokens":813,"num_sources_used":0,"completion_tokens":93,"cost_in_usd_ticks":40288000,"prompt_tokens_details":{"text_tokens":813,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":380,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":813,"tokens_out":93,"duration_ms":8196,"temperature":1.0,"reasoning_tokens":380,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T20:25:51.921697+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain the five LIBERO+ tasks with the identical LLM pipeline but replace the generated task spaces by random or empty parameterizations; if success rates then collapse to the near-zero levels of the sparse-reward and no-ACL baselines, the claim that the LLM-generated specifications are doing the essential work is falsified.","supporting_citations":[],"review_version":1}