{"id":"a46ab21f-1b1f-44db-b52f-3be48920e52d","arxiv_id":"2607.04591","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Simple-to-complex staged demonstration collection (task decomposition, environment standardization, progressive complexity) substantially improves π0.5 VLA success on dual-arm block sorting and towel folding versus end-to-end demos.","lead":"Organizing robot demos from simple skills to full tasks (S2C) raised dual-arm VLA success from 0% to 80% on block sorting and 0% to 25% on towel folding versus end-to-end trajectories. This suggests data structure, not only model scale, can unlock long-horizon manipulation under limited demos.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Unmatched demo counts and stage standardization confound the causal claim that S2C organization itself drives the success-rate gains.","rationale":"The reader correctly isolates the weakest assumption: unequal demonstration counts plus stage-specific scene standardization that alter the training distribution. The paper is explicit that counts differ and that the comparison is not matched-N (§5.1, Table 3). Evaluation N is also tiny (5 block trials), there are no ablations of the three principles, no seeds/error bars, and only one model. These factors leave the causal claim under-supported, so CONDITIONAL remains the right verdict; the stress test does not move it. A single matched-budget re-run would settle whether organization, rather than extra data or easier intermediate distributions, is doing the work.","tokens_in":14650,"tokens_out":532,"duration_ms":5754,"concrete_test":"Collect a matched-N direct baseline of 300 complete end-to-end trajectories (same total budget as S2C) under the same platform/cameras, retrain π0.5 with identical hyperparameters, and re-evaluate on the same 5+28 held-out trials. If success remains near 0% while S2C stays high, the organization claim is supported; if the matched direct baseline rises substantially, the gains are largely budget/distribution artifacts.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Table 3, §5) is that S2C organization—task decomposition, environment standardization, progressive complexity—causes the jump from 0% to 80%/25% under identical π0.5 training. The paper itself states that demo counts differ (300 staged vs 200 direct) and that results evaluate the strategy as a whole rather than a matched-N comparison (§5.1). Stage-wise standardization further changes the training distribution: Stage 1 uses single-color scenes and independent arms; Stage 2 multi-color but still single-arm; only Stage 3 is dual-arm full task. Direct collection uses full dual-arm trajectories from the start with higher initial-state diversity (especially towels). Thus the observed gains could be produced by (i) 50% more data, (ii) easier intermediate distributions that never appear in the direct baseline, or (iii) the curriculum itself. Without a matched-budget or ablated control, the causal attribution to the three S2C principles remains unsecured.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that demonstration organization is an underexplored but load-bearing factor in Vision-Language-Action (VLA) imitation learning. It proposes a simple-to-complex (S2C) collection strategy with three principles—capability-level task decomposition (basic manipulation → object/state understanding → task execution), environment standardization within stages, and progressive complexity scheduling—and instantiates it on a dual-arm SO-101 platform for block grasping-and-sorting and towel folding. Under a fixed π0.5 model, training pipeline, and evaluation protocol, S2C is compared to conventional end-to-end full-task trajectories. Table 3 reports large success-rate gains: 80% (4/5) vs 0% (0/5) on blocks and 25% (7/28) vs 0% (0/28) on towels, with qualitative failure analysis attributing residual errors mainly to low-cost hardware and deformable-state variability rather than semantic misunderstanding.","tokens_in":14976,"tokens_out":1240,"duration_ms":14735,"significance":"If the causal claim holds, the work would usefully reorient limited-data VLA practice from architecture and scale alone toward how demonstrations are structured, with a reusable capability progression that covers both rigid and deformable dual-arm tasks. Strengths include a controlled same-model/same-pipeline design that isolates collection strategy as the intended variable, real-robot dual-arm evaluation, explicit stage protocols (§4.3.1–4.3.2), and honest failure analysis (§5.2). The contribution is primarily empirical and methodological rather than theoretical; its impact depends on whether gains can be attributed to S2C organization rather than unmatched data volume or easier intermediate distributions. With stronger controls, this would be a practical and timely addition to robotic imitation-learning literature.","major_comments":[{"comment":"§5.1 and Table 3: the central causal claim—that S2C organization (decomposition, standardization, progressive complexity) drives the jump from 0% to 80%/25%—is not secured by the experimental design. S2C uses 300 demonstrations vs 200 for the direct baseline, and the paper states results evaluate the strategy as a whole rather than a matched-N comparison. A matched-budget control (e.g., 300 direct full-task demos, or 200 S2C demos) is needed; without it, gains may simply reflect 50% more data.","section":null},{"comment":"§4.3 and §5: stage-wise environment standardization and single-arm intermediate stages change the training distribution relative to the direct baseline (full dual-arm trajectories with higher initial-state diversity, especially for towels). There are no ablations that isolate the three stated principles—e.g., decomposition without standardization, matched full-task scenes with curriculum ordering only, or staged data mixed without progressive scheduling. As written, one cannot tell which principle, if any, is necessary for the reported gains.","section":null},{"comment":"Table 3, block task: evaluation uses only 5 trials (4/5 vs 0/5). With such a small N_eval and no multi-seed training or confidence intervals, the 80% figure is too fragile to support a strong comparative claim. At minimum, report substantially more held-out trials, multiple random seeds or teleop operators, and uncertainty estimates; the towel N=28 is better but still lacks statistical testing against the 0% baseline.","section":null},{"comment":"Abstract and §1 claim improvements in “training stability” as well as success rate, but §5 reports only final success rates and qualitative failure modes. No learning curves, loss/variance metrics, or intermediate checkpoint evaluations are provided. Either quantify stability (e.g., success vs training steps across seeds) or narrow the claim to task success under the reported protocol.","section":null}],"minor_comments":[{"comment":"ACM front matter still contains placeholders (“Make sure to enter the correct conference title…”, “https://doi.org/XXXXXXX.XXXXXXX”); clean for submission.","section":null},{"comment":"§2.3 correctly notes that curriculum learning usually structures training rather than collection; a tighter comparison to hierarchical imitation learning, skill libraries, and staged teleoperation protocols would better situate novelty.","section":null},{"comment":"Table 1 and §4 use slightly different stage labels for towels (initial-state handling / state normalization / rule-based folding vs D_manip / D_state / D_exec); align terminology for readability.","section":null},{"comment":"Figures 3–8 are useful stage illustrations but lack quantitative scene statistics (object pose ranges, color layouts, towel configuration diversity); adding these would make standardization reproducible.","section":null},{"comment":"§5.2 failure ratios are conditioned on failure events, not trials; state this more prominently so readers do not misread 50%/90% as trial-level failure probabilities.","section":null},{"comment":"References [23] and [24] are the authors’ concurrent arXiv preprints used for model choice and future-work framing; ensure they are cited only where necessary and that model selection is justified independently of those works.","section":null}],"recommendation":"major_revision","confidential_remarks":"The empirical signal is directionally interesting and the same-pipeline design is a real strength, but the manuscript currently overclaims causality relative to the evidence. I would accept after major revision if the authors add matched-budget and/or ablation controls and strengthen evaluation statistics; without those, the paper is closer to a useful systems note than a conclusive result on demonstration organization. Scope is appropriate for a robotics/learning journal if the experimental bar is raised."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing worth knowing: under the same π0.5 model and pipeline, their staged collection protocol gets real dual-arm success (80% on 5 block trials, 25% on 28 towel trials) where direct full-task demos get zero. That is a usable empirical signal for people collecting limited real-robot VLA data, not a foundational result.\n\nWhat is actually new is the concrete three-stage collection recipe for two hard dual-arm tasks (rigid color sorting and deformable folding), with explicit capability ordering (basic manip → state understanding → full execution), environment standardization, and progressive complexity. They isolate the collection strategy as the only variable, document the failure modes of the direct baseline clearly (arm asymmetry, post-grasp hesitation), and give enough stage definitions that someone could try to reproduce the protocol. Related-work placement is honest: they cite curriculum and hierarchical RL and correctly position themselves as applying that idea to modern VLA fine-tuning under limited data rather than inventing curriculum.\n\nThe soft spots are real but proportional. Demo counts are unmatched (300 staged vs 200 direct) and the paper says so; stage standardization also changes the training distribution (single-color/single-arm early, dual-arm only late). So the jump from 0% cannot yet be cleanly attributed to the three S2C principles versus more data or easier intermediate distributions. Eval N is tiny for blocks, no seeds/error bars/ablations, single model, no public code or data. Those are standard limited-data robotics weaknesses, not hidden math or circularity problems; the evaluation is straightforward held-out success rates.\n\nThis is for people who actually collect dual-arm or deformable demos and care about data efficiency more than architecture. A serious editor should send it to referees; the claim is important enough in that niche and the design is readable enough to deserve a proper review, even if the revision will demand matched-budget controls and broader validation. I would read the protocol if I were collecting similar data; I would not cite the causal claim until the confounds are closed.","headline":"Practical dual-arm VLA recipe that turns 0% end-to-end demos into non-zero success, but the causal claim for S2C is confounded by unmatched data budgets and stage-wise distribution shifts.","tokens_in":15584,"tokens_out":531,"would_cite":false,"duration_ms":5530,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"How you organize robot demos matters as much as model architecture for long-horizon VLA learning.","keywords":["Vision-Language-Action","imitation learning","demonstration collection","curriculum learning","dual-arm manipulation","long-horizon manipulation","deformable objects","robot learning"],"falsifier":"Re-run the same π0.5 fine-tuning and evaluation with matched total demonstration counts and comparable scene diversity between S2C and direct end-to-end collection; if the large success-rate gap (80% vs 0% blocks; 25% vs 0% towels) disappears, the organization claim is undermined.","tokens_in":15495,"feed_emoji":"🤖","tokens_out":636,"duration_ms":5092,"temperature":0.7,"pith_summary":"This paper argues that Vision-Language-Action models for robots have focused almost entirely on architectures, training recipes, and dataset size, while neglecting how human demonstrations are collected and ordered. The authors claim that naively recording full end-to-end trajectories for long-horizon tasks produces highly coupled, heterogeneous data that is hard to learn from under limited data. They instead propose a simple-to-complex collection strategy: decompose a task into ordered capability stages (basic manipulation, then object/state understanding, then full execution), standardize the scene within each stage to remove irrelevant variation, and only then raise complexity. On a dual-arm platform they show that policies trained this way succeed far more often than policies trained on the same model and pipeline with conventional complete-task demos. The practical stake is that dataset construction itself can become a controllable lever for skill acquisition and long-horizon reliability, not merely a quantity problem.","feed_headline":"How robot demos are ordered can decide if policies work","feed_subtitle":"Simple-to-complex stages beat full end-to-end trajectories on dual-arm sorting and towel folding","key_machinery":"Simple-to-Complex (S2C) structured demonstration collection: a three-principle process that decomposes a long-horizon task into ordered capability stages (basic manipulation, object perception and state understanding, task execution), standardizes the environment within each stage, and schedules progressive complexity so the policy acquires prerequisite skills before full composition.","core_discovery":"Under identical model, training, and evaluation conditions, a simple-to-complex structured demonstration collection strategy—task decomposition into progressive sub-skills, environment standardization, and increasing complexity—yields substantially higher real-robot success rates than directly collecting complete end-to-end task trajectories for both rigid-object block sorting and deformable towel folding.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Simple-to-complex demos beat end-to-end trajectories for VLA policies","Organizing robot demos by rising complexity raises success rates","Structured sub-skill demos outperform full trajectories on dual-arm tasks","Progressive demo order lifts sorting and folding success under same model","Demonstration structure drives VLA learning efficiency and stability"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The gains are caused by the S2C organization principles themselves, not by the unequal demonstration counts (300 staged versus 200 direct) or by stage-specific scene choices that simply change the training distribution.","fun_headline_variants_meta":{"raw":{"variants":["Simple-to-complex demos beat end-to-end trajectories for VLA policies","Organizing robot demos by rising complexity raises success rates","Structured sub-skill demos outperform full trajectories on dual-arm tasks","Progressive demo order lifts sorting and folding success under same model","Demonstration structure drives VLA learning efficiency and stability"]},"model":"grok-4.5","effort":"low","cost_usd":0.005582,"raw_usage":{"total_tokens":1503,"prompt_tokens":811,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":55820000,"prompt_tokens_details":{"text_tokens":811,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":619,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":811,"tokens_out":73,"duration_ms":5185,"temperature":1.0,"reasoning_tokens":619,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T16:45:06.086551+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same π0.5 fine-tuning and evaluation with matched total demonstration counts and comparable scene diversity between S2C and direct end-to-end collection; if the large success-rate gap (80% vs 0% blocks; 25% vs 0% towels) disappears, the organization claim is undermined.","supporting_citations":[],"review_version":1}