{"id":"f9b6b246-e0d2-44c4-9eaf-ce28df114e77","arxiv_id":"2510.18608","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The paper proposes continual learning plus compositional adapters as the path to adaptable robot foundation models, but its two supporting experiments are underspecified and reliant on the authors' earlier work.","lead":"This short paper argues that AI systems should be built by composing small, continually-learned modules instead of retraining giant foundation models, and reports two preliminary experiments meant to support that view. A generalist might read it as a snapshot of an ongoing agenda for making robots adapt to new tasks without full retraining.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II's robot comparison is not interpretable: environment unnamed, compute budgets unequal, and baselines likely under-provisioned, so the headline win (0.91 vs 0.0) may be a resource artifact.","rationale":"The reader's weakest assumption identified the fairness of the robot comparison under equal compute; my analysis agrees and sharpens it: the table's training times contradict the 'same computational constraints' claim, making the comparison internally inconsistent. This is the single most load-bearing issue because the paper's headline experimental win (0.91 vs 0.0 success) and the phrase 'we show' depend on it. If resolved, the verdict would still remain conditional until the environment is named, seeds/repeated runs are provided, and the image-classification results are confirmed not merely to reproduce prior work. Other concerns, such as the §III diminishing-returns assertion based on a contested preprint, are less load-bearing because they support motivation rather than the central empirical claim. Thus, the reader's CONDITIONAL verdict is appropriate and no change is needed.","tokens_in":3534,"tokens_out":2847,"duration_ms":26819,"concrete_test":"On a named, public manipulation benchmark (e.g., Meta-World MT1 or ManiSkill3), train OpenVLA and WSA with identical demonstrations, seeds, and hardware. Record success rate versus training time. Run OpenVLA to its published convergence schedule (even if >92h) and also at a compute-matched budget. If OpenVLA's converged success rate meets or exceeds WSA's 0.91, the paper's claim of superiority is an artifact of under-training; if OpenVLA remains near 0 even when properly converged, the claim is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that composing models yields strong performance with limited resources rests mainly on Table II. The text (Sec. IV) says WSA beats OpenVLA/InstructRL 'given the same computational constraints,' but the table reports training times of 14h vs 40h/92h — evidently not the same budget. The environment, observation space, action space, dataset size, and number of seeds are all unspecified, and no error bars are given for the robot results. OpenVLA is a 7B VLA that typically requires large-scale data and long training schedules; reporting a 0.0 success rate after 92h (without specifying hardware parallelization) is consistent with a model that has not been fine-tuned enough, i.e., resource under-provisioning, not necessarily inferiority. The success-rate metric without seeds cannot be assessed. If this comparison is removed or corrected, the central claim of strong robot performance rests on Table I alone, which is itself drawn from the authors' prior work [5] without disclosure. Thus, the experimental support for the paper's main claim is currently unverifiable and likely confounded.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a compositional paradigm for foundation models, combining continual learning and compositionality to adapt to dynamic tasks. It presents two sets of experiments: (1) CUB200 image classification with hierarchical merging of LoRA adapters (HAM), reporting 55.17% accuracy vs 47.56% for SD-LoRA; (2) robotic manipulation with a 'WSA' architecture that reportedly achieves 0.91 success rate versus 0.0 for OpenVLA and InstructRL under apparently lower training time. The conclusion claims that composing multiple models enables strong performance with limited resources.","tokens_in":3712,"tokens_out":3511,"duration_ms":29626,"significance":"If the results held, the paper would support a promising and cost-efficient alternative to monolithic scaling for adapting foundation models, with implications for robotics. The compositional continual-learning direction is timely. However, the experimental support is currently not independently verifiable: the robot comparison is underspecified and likely confounded by unequal resources; the image experiments are described too briefly; and both sets of results appear to be drawn from the authors' own prior papers without appropriate disclosure or artifact release. The paper would be valuable as a research proposal or manifesto, but as a technical paper it needs substantial revision.","major_comments":[{"comment":"The robotic manipulation comparison is not interpretable. The task, environment, observation and action spaces, dataset size, number of demos, and evaluation protocol are not named. 'Reward/Step' is not defined, and no error bars or seeds are given. Critically, the text claims the baselines fail 'given the same computational constraints,' but Table II reports 14 h vs 40 h and 92 h training times, with hardware unspecified; this is consistent with resource under-provisioning of OpenVLA/InstructRL. The central success claim (0.91 vs 0.0) therefore rests on an unverifiable and possibly confounded comparison. Please specify exact compute, training steps, and provide per-seed results and learning curves.","section":"Section IV, Table II"},{"comment":"The image-classification claim is underspecified: '50 tasks on CUB200' does not state how the 200 classes were split (e.g., 2-class incremental tasks), the task order, or whether this is class-incremental or task-incremental evaluation. Only mean accuracy is reported; no per-task accuracy, forgetting, or task-order ablation is given. Without these details, the 55.17 vs 47.56 advantage cannot be assessed as a continual-learning result. Provide the full protocol and, ideally, release code.","section":"Section IV, Table I"},{"comment":"The experimental evidence is not original to this submission but is taken from the authors' prior works [5] and [12], with the framing from [3]; none of this is disclosed in Section IV. This creates a circularity/verifiability problem: a reader cannot separate the claimed results from self-citation. Please state explicitly which results are new, and either include full descriptions from the prior papers or provide independent re-runs with artifacts.","section":"Section IV (attribution)"},{"comment":"The premise that scaling has hit diminishing returns is asserted from a single preprint [10] and presented as fact without engaging counter-evidence. Since the paper's motivation is that compositionality outperforms scaling, this load-bearing assumption needs stronger support or a more nuanced statement.","section":"Section III"}],"minor_comments":[{"comment":"Reference [1] does not seem to be about foundation models generally (it is 'Learning interactive real-world simulators'); please check relevance.","section":"Abstract/Introduction"},{"comment":"Typo: 'we propose a alternative vision' should be 'we propose an alternative vision'.","section":"Section I"},{"comment":"Typo: 'where its clear' should be 'where it is clear'.","section":"Section IV"},{"comment":"There is no related-work section positioning the paper with respect to other compositional/continual learning methods; consider adding one.","section":"General"},{"comment":"Define 'Reward/Step' and specify how success rate is measured and over how many episodes.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"This manuscript reads as a compact summary of the authors' prior publications (HAM and WSA) rather than a self-contained technical paper. The robot comparison, as presented, is not scientifically defensible because of missing environment details and unequal compute budgets. I would ask the editor to require a much more detailed experimental section or to treat the paper as a position/outlook piece rather than a results paper. Note also that refs [3], [5], and [12] are all self-citations used as the basis for the main claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a short position paper arguing that continual learning and compositionality — hierarchical merging of LoRA adapters and a small composed agent — should replace monolithic scaling for foundation models, especially in robotics. The direction is reasonable, and the paper is readable. But as a standalone manuscript, its experimental support is too weak to carry the claims.\n\nWhat's actually new: not much. The framing comes from the authors' own position paper [3]; the two experiments are re-presentations of their prior work (HAM in [5], WSA in [12]). The text adds no new method, dataset, or derivation. That's not a crime if disclosed, but here the tables appear as \"our results\" without saying they're drawn from earlier papers.\n\nWhat it does well: it names a real problem — foundation models are expensive to adapt, and catastrophic forgetting matters in robotics — and it points to a plausible toolbox (LoRA composition, small pretrained components). The image classification experiment at least has a dataset (CUB200), error bars, and training times. If HAM genuinely beats SD-LoRA and InfLoRA at lower cost, that's a useful incremental result.\n\nThe soft spots are load-bearing. Table II, the headline robot result, is not interpretable. The environment is never named; no observation or action space, no seeds, no error bars. The caption says \"reduced training times,\" but the text claims the comparison is \"given the same computational constraints\" while the training times are 14h vs 40h and 92h. That's not the same budget. OpenVLA at 92h with 0.0 success is consistent with an under-provisioned baseline, not evidence that composition is better. A 0.91 success rate without variance or task definition means nothing. Also, Table I says \"50 tasks on CUB200\" without specifying the class split, order, or stream setup — details that matter for continual learning.\n\nThe diminishing-returns premise in Section III rests on a single contested preprint [10], with no counter-evidence. The conclusion \"we show that composing multiple models enables strong performance\" overstates what this manuscript alone demonstrates.\n\nIf the authors fixed these issues — name the robot task, equalize compute or justify why unequal budgets are fair, add seeds and per-task curves, and state plainly that the experiments come from prior work — this could be an acceptable short position paper. As it stands, I wouldn't trust Table II.\n\nWho is this for? Someone wanting a quick, readable pitch for compositionality in foundation-model robotics. It isn't citable yet. I'd read the referenced prior work before recommending anything. Send it to peer review only as a position/workshop paper, and only if the robot comparison is corrected or cut. Otherwise, desk reject.","headline":"The direction is sensible, but the evidence is almost entirely self-cited prior work, and the robot win is uninterpretable as reported.","tokens_in":4367,"tokens_out":3640,"would_cite":false,"duration_ms":31575,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that composing small, specialized models—rather than scaling one monolithic foundation model—lets agents match or beat much larger systems in both vision and robotic manipulation while using a fraction of the training budg","keywords":["foundation models","continual learning","compositionality","LoRA adapters","hierarchical adapter merging","robotic manipulation","vision-language-action models","parameter-efficient fine-tuning"],"falsifier":"Train OpenVLA or InstructRL on the same, named manipulation task with their intended full compute budgets (not capped at 14 hours) and measure success rate; if either exceeds 0.91, the composed agent's win is a resource-starvation artifact. Alternatively, run WSA on a standard multi-task manipulation benchmark and compare against published baseline success rates.","tokens_in":3320,"feed_emoji":"🤖","tokens_out":10637,"duration_ms":83425,"temperature":0.7,"pith_summary":"The paper aims to establish that foundation models do not have to be retrained or scaled to keep learning: composing multiple small, specialized modules can yield stronger performance with far less compute. In image classification on CUB200 with 50 tasks, a hierarchical adapter merging method reaches 55.17% accuracy, beating SD-LoRA (47.56%) and InfLoRA (36.02%) while training faster. In robotic manipulation, a small composed agent achieves a 0.91 success rate after 14 hours of training, compared with 0.0 for OpenVLA and InstructRL when those baselines are given the same computational budget. The authors read these results as evidence that compositionality and continual learning are a viable alternative to the scaling-only paradigm for developing adaptive, efficient AI.","feed_headline":"0.91 vs 0.0: small composed agent beats OpenVLA","feed_subtitle":"The same compositional recipe lifts CUB200 accuracy to 55.17% and trains a robot agent in 14 hours.","key_machinery":"Two mechanisms carry the argument. The first is hierarchical adapter merging (HAM): for each task, a separate LoRA adapter—a small set of low-rank trainable matrices—is trained, then merged with previously learned adapters first within similarity groups and then across groups into a single module; this ordering is what the paper says reduces merging interference. The second is WSA, a lightweight robotic architecture that combines small off-the-shelf pre-trained models with small adapters and an attention mechanism to dynamically scale pre-trained components to the current context. Both are instances of the same thesis: compose rather than scale.","core_discovery":"On the paper's own terms, the discovery is that compositionality rescues foundation models from task rigidity and prohibitive retraining cost. The authors demonstrate this in two settings: a vision system that trains a separate LoRA adapter for each task and merges them hierarchically, first within similarity groups and then into one module, reducing interference and improving accuracy; and a robotic manipulation agent built from small pre-trained components, small adapters, and an attention mechanism that learns a manipulation skill in 14 hours with a 0.91 success rate, while OpenVLA and InstructRL under identical compute constraints score 0.0. The unifying claim is that a modular, continua","pith_inferences":["The paper leaves the manipulation environment unnamed and does not specify the action or observation space; until those are disclosed, the 0.91 vs 0.0 success gap cannot be independently reproduced, so the obvious next step is to run WSA on a public multi-task manipulation benchmark with adequately provisioned baselines.","The motivation that scaling has hit diminishing returns is borrowed from a single contested preprint; even if that premise is wrong, the compositional mechanism could still be useful as a complementary strategy, a possibility the paper does not discuss.","If adapter merging truly resists interference, a natural testable extension is to apply HAM to large language models on sequential instruction-tuning tasks and measure forgetting against existing continual-learning baselines.","The paper implicitly assumes a single composed model must solve every seen task; an alternative design that routes each input to the relevant adapter might trade a little memory for even less interference, but that route is not explored."],"forward_implications":["If HAM holds up, adapting a foundation model to a new task becomes as cheap as training one small adapter and merging it, enabling continual learning without architectural changes or full retraining.","If WSA's robotic result is accurate, robot skill acquisition could drop from days of GPU time to hours, making learning-based manipulation practical for smaller research groups and edge robots.","The hierarchical merging order specifically reduces interference, implying that merge order matters and can be chosen to preserve performance across unrelated tasks.","The same compositional design should extend to any modality with a frozen backbone and small trainable adapters, since it does not rely on task-specific architecture."],"fun_headline_variants":["Compositional agent scores 0.91, OpenVLA 0.0 in 14 hours","Small composed robot beats OpenVLA: 0.91 vs 0.0","Compositionality rescues AI: 0.91 success, OpenVLA zero","Modular agent trains in 14h, outperforms OpenVLA 0.91-0.0","Tiny LoRA agents beat OpenVLA with compositional merging"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The robotic result assumes that giving OpenVLA and InstructRL the same tiny training budget as the small agent is a fair comparison, and that the unnamed robot task is a meaningful test; if either is not true, the 0.91 vs 0.0 success gap does not prove the composed agent is better.","fun_headline_variants_meta":{"raw":{"variants":["Compositional agent scores 0.91, OpenVLA 0.0 in 14 hours","Small composed robot beats OpenVLA: 0.91 vs 0.0","Compositionality rescues AI: 0.91 success, OpenVLA zero","Modular agent trains in 14h, outperforms OpenVLA 0.91-0.0","Tiny LoRA agents beat OpenVLA with compositional merging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00052,"raw_usage":{"total_tokens":2285,"prompt_tokens":608,"completion_tokens":1677,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":352,"completion_tokens_details":{"reasoning_tokens":1562}},"tokens_in":352,"tokens_out":1677,"duration_ms":10823,"temperature":1.0,"reasoning_tokens":1562,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:49:26.251520+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train OpenVLA or InstructRL on the same, named manipulation task with their intended full compute budgets (not capped at 14 hours) and measure success rate; if either exceeds 0.91, the composed agent's win is a resource-starvation artifact. Alternatively, run WSA on a standard multi-task manipulation benchmark and compare against published baseline success rates.","supporting_citations":[],"review_version":1}