{"id":"907fe536-cfbc-424e-9b6b-14f16246c92f","arxiv_id":"2605.16056","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper introduces health-conditioned VLA models that incorporate a health vector via a new projector module and train on 128 malfunction episodes in the LIBERO simulator to complete tasks despite degraded joints.","lead":"This paper adds a Health Projector module to vision-language-action models so the robot can use a health vector describing joint angles and torques to adapt its actions when parts degrade. A smart generalist might read it to see a practical way to keep robots working longer in real settings instead of stopping at the first mechanical issue.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Generalization claim rests on 128 simulated episodes covering representative degradations","rationale":"The reader's weakest assumption directly identifies the data representativeness and health-vector fidelity as the critical unverified conditions for the adaptation claim. This matches the load-bearing point: without evidence that the 128 episodes suffice for generalization beyond the training distribution, the headline result remains provisional. Adjusting to CONDITIONAL reflects that the architecture addition is lightweight and plausible, but the empirical support needs the proposed check before stronger acceptance.","tokens_in":1725,"tokens_out":336,"duration_ms":29129,"concrete_test":"Partition the 128 episodes into train/test splits with held-out degradation configurations (e.g., different joints or severity levels); retrain the projector on the train split only and evaluate success rate on the held-out test episodes versus the baseline VLA-Adapter. If the health-conditioned model does not maintain a clear advantage on the unseen patterns, the generalization assumption does not hold.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that injecting a lightweight Health Projector allows the VLA to succeed on Libero-Spatial tasks with varied joint degradations where the baseline VLA-Adapter fails. This requires that the 128 teleoperated malfunction episodes span a sufficiently diverse and representative set of degradation patterns (different joints, angle limits, torque reductions) and that the supplied health vector accurately encodes current capabilities at inference time. If the collected episodes are narrow in scope or the vector is an oracle signal rather than estimated from real sensor data, success on the reported configurations would not establish adaptation to unseen or real-world degradations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes malfunction-aware Vision-Language-Action (VLA) models by augmenting VLA-Adapter with a lightweight Health Projector module. This module conditions the model on a health vector that encodes joint operation angles and torque capabilities, enabling adaptation to physical degradations such as joint failures. The approach is demonstrated by collecting and training on 128 teleoperated malfunction episodes in the LIBERO simulation environment for Libero-Spatial tasks, with the claim that the resulting model succeeds on degraded configurations where the default pretrained VLA-Adapter fails.","tokens_in":1846,"tokens_out":590,"duration_ms":54927,"significance":"If the empirical results hold and generalize, the work addresses a practical gap in deploying VLA models on physical robots subject to hardware degradation. The lightweight projector design avoids full retraining and could support efficient adaptation; releasing the code and dataset would further strengthen reproducibility and enable follow-on research in robust robot control.","major_comments":[{"comment":"Data Collection section: the central generalization claim depends on the 128 teleoperated episodes spanning a representative range of degradation patterns (different joints, angle limits, torque reductions). The manuscript must report the exact distribution of these patterns and any held-out test configurations to establish that success is not limited to the collected set.","section":"Data Collection"},{"comment":"Results section: the abstract states successful operation but supplies no quantitative success rates, baseline comparisons (e.g., against fine-tuned VLA-Adapter or oracle health signals), ablation studies on the projector, or error analysis. These metrics are load-bearing for the claim that the health-conditioned model outperforms the default pretrained model.","section":"Results"},{"comment":"Health vector definition and inference: the paper must clarify whether the health vector is an oracle signal supplied during training and evaluation or estimated from real sensor data at deployment time; if the former, the adaptation claim does not yet extend to realistic malfunction detection.","section":"Methods / Health Projector"}],"minor_comments":[{"comment":"Abstract: the citation to LIBERO [1] should be expanded to a full reference; the phrase 'different configurations of degraded joints' would benefit from a brief parenthetical example of the degradation types tested.","section":"Abstract"},{"comment":"Notation: ensure consistent use of 'health vector' versus 'Health Projector' throughout; a diagram showing the exact insertion point of the projector into the VLA-Adapter would improve clarity.","section":"Architecture"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads as an early empirical exploration; the journal's typical standards for robotics papers emphasize extensive real-robot validation and statistical rigor, so the authors may need to expand the evaluation substantially before acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on our manuscript. The comments highlight important areas for improving clarity, rigor, and completeness, particularly regarding data details, quantitative evaluation, and methodological clarifications. We address each major comment below and will incorporate revisions to strengthen the paper.","responses":[{"response":"We agree that reporting the distribution of degradation patterns is necessary to support the generalization claims. In the revised manuscript, we will expand the Data Collection section with a table detailing the breakdown of the 128 episodes by joint type, angle limit reductions, and torque capability decreases. We will also explicitly describe the held-out test configurations and how they differ from the training degradations to demonstrate that performance is not limited to the collected set.","revision_made":"yes","referee_comment":"[Data Collection] Data Collection section: the central generalization claim depends on the 128 teleoperated episodes spanning a representative range of degradation patterns (different joints, angle limits, torque reductions). The manuscript must report the exact distribution of these patterns and any held-out test configurations to establish that success is not limited to the collected set."},{"response":"We acknowledge that the current version lacks the quantitative metrics needed to fully substantiate the performance claims. We will revise the Results section to include success rates for the health-conditioned model compared to the baseline pretrained VLA-Adapter, additional baselines such as fine-tuned VLA-Adapter and oracle health signals, ablation studies on the Health Projector module, and an error analysis of failure modes. These additions will provide a more complete and rigorous evaluation of the approach.","revision_made":"yes","referee_comment":"[Results] Results section: the abstract states successful operation but supplies no quantitative success rates, baseline comparisons (e.g., against fine-tuned VLA-Adapter or oracle health signals), ablation studies on the projector, or error analysis. These metrics are load-bearing for the claim that the health-conditioned model outperforms the default pretrained model."},{"response":"We appreciate this clarification request. In the present work, the health vector is supplied as an oracle signal during both training and evaluation in the LIBERO simulation. We will update the Methods section to state this explicitly. We will also add a limitations paragraph and future work discussion on estimating the health vector from onboard sensor data to extend the approach toward realistic deployment scenarios.","revision_made":"yes","referee_comment":"[Methods / Health Projector] Health vector definition and inference: the paper must clarify whether the health vector is an oracle signal supplied during training and evaluation or estimated from real sensor data at deployment time; if the former, the adaptation claim does not yet extend to realistic malfunction detection."}],"tokens_in":1411,"tokens_out":578,"duration_ms":55083,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that the authors inject a Health Projector module into VLA-Adapter so the policy can adjust for degraded joints using an explicit health vector of angles and torques. They train on 128 teleoperated malfunction episodes in the Libero-Spatial setup and report that the modified model completes tasks the baseline cannot handle with the same degradations.","headline":"The paper adds a Health Projector to condition VLA models on joint health vectors, but the 128 simulated episodes leave the generalization story thin.","tokens_in":2323,"tokens_out":147,"would_cite":false,"duration_ms":36748,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/AbsoluteFloorClosure.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"We propose a new policy that is health-conditioned: ât:t+C=πθ(ot,l,h) where the health vector h is projected into the model’s latent space via an MLP and fused with the action prediction head"},{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"Health Projector: two-layer MLP … fh=W2 GELU(W1 h+b1)+b2 … zero-initialized so that fh=0 at the start of training"}],"headline":"Health-conditioned VLA projector for joint degradation has no structural overlap with RS cost or distinction forcing","alignment":"orthogonal","rationale":"The paper's central machinery is a lightweight two-layer MLP Health Projector that injects a 7D health vector h∈[0,1]^7 (joint angle/torque limits) into the proprioceptive tokens of a frozen VLA-Adapter action head, trained on 128 simulated malfunction episodes. This is a standard parameter-efficient fine-tuning technique for fault-tolerant robot control. It invokes none of the RS primitives: no reciprocal cost J(x)=½(x+x⁻¹)−1, no φ-ladder, no 8-tick periodicity, no Alexander-duality forcing of D=3, and no parameter-free derivation of constants. The domain (VLA adaptation on LIBERO Spatial) lies outside the RS forcing chain.","tokens_in":45832,"confidence":"high","tokens_out":385,"duration_ms":28881,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A lightweight health projector module added to vision-language-action models lets robots adapt to degraded joints and finish tasks where standard models fail.","keywords":["vision-language-action models","malfunction-aware control","health-conditioned VLA","robot joint degradation","LIBERO environment","VLA-Adapter","health projector","teleoperated episodes"],"falsifier":"Running the health-conditioned model on joint degradation patterns that differ from those in the 128 training episodes and finding that it fails at the same rate as the unmodified baseline, or observing that inaccurate health vectors cause the model to produce ineffective actions, would falsify the central claim.","tokens_in":2615,"feed_emoji":"🤖","tokens_out":707,"duration_ms":68700,"temperature":0.7,"pith_summary":"The paper shows how to make vision-language-action models aware of a robot's physical condition by feeding them a health vector that describes joint angles and torque limits. The authors add a small Health Projector module to the existing VLA-Adapter and train it on 128 teleoperated episodes of simulated joint malfunctions collected in the LIBERO environment. This matters because everyday robots suffer gradual wear that causes current systems to stop working on assigned tasks. If the approach holds, robots could keep operating through partial hardware failures instead of needing immediate fixes or full retraining. The reported outcome is that the modified model succeeds across different degradation setups on spatial tasks while the unmodified pretrained version cannot.","feed_headline":"Health projector lets VLA robots handle degraded joints","feed_subtitle":"Lightweight module conditions action predictions on joint angles and torque so spatial tasks still succeed after physical wear.","key_machinery":"The Health Projector module, which accepts a health vector of joint operation angles and torque capabilities and conditions the model's action predictions to account for physical degradation.","core_discovery":"By injecting a Health Projector module into the VLA-Adapter architecture and training it on a dataset of 128 teleoperated malfunction episodes collected in the LIBERO environment, the health-conditioned model can successfully complete spatial tasks using degraded joints, whereas the unmodified Libero-Spatial-Pro model cannot.","pith_inferences":["The same conditioning approach could be tested with online estimation of the health vector from onboard sensors instead of an external input.","The method might extend to additional failure modes such as gripper weakness or sensor drift beyond joint degradation.","Direct transfer experiments from the simulation-trained model to physical robots would clarify how well the learned adaptations hold in real hardware.","Combining this health conditioning with predictive maintenance alerts could reduce downtime in deployed robot systems."],"forward_implications":["The model can adjust its behavior to varied configurations of degraded joints without retraining the entire pretrained VLA-Adapter.","Task success becomes possible even when joint angles or torque outputs are reduced below nominal levels.","Only a small module addition is required rather than a full redesign of the vision-language-action pipeline.","The trained adaptation generalizes across different degradation patterns encountered in the collected simulation data."],"fun_headline_variants":["Health data conditions VLA for degraded robot joints","VLA adapts to faulty joints with health vector input","Health projector trains VLA on malfunction episodes","Health conditioning helps VLA handle joint degradation"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The health vector supplied to the projector accurately captures the robot's current joint operation angles and torque capabilities, and the 128 teleoperated malfunction episodes collected in simulation are representative enough for the model to generalize across varied degradation patterns.","fun_headline_variants_meta":{"raw":{"variants":["Health data conditions VLA for degraded robot joints","VLA adapts to faulty joints with health vector input","Health projector trains VLA on malfunction episodes","Health conditioning helps VLA handle joint degradation"]},"model":"grok-4.3","cost_usd":0.007972,"raw_usage":{"total_tokens":3534,"prompt_tokens":637,"num_sources_used":0,"completion_tokens":56,"cost_in_usd_ticks":79715500,"prompt_tokens_details":{"text_tokens":637,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2841,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":637,"tokens_out":56,"duration_ms":49330,"temperature":1.0,"reasoning_tokens":2841,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-20T18:13:00.972366+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Running the health-conditioned model on joint degradation patterns that differ from those in the 128 training episodes and finding that it fails at the same rate as the unmodified baseline, or observing that inaccurate health vectors cause the model to produce ineffective actions, would falsify the central claim.","supporting_citations":[],"review_version":1}