{"id":"37daf5f9-cd25-415d-8673-b7667933ef23","arxiv_id":"2606.25374","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new settled-gate benchmark shows hybrid estimate-then-control with recurrent gain inference reaches 94-98% success on unseen sign/gain actuator faults while end-to-end RL and PD score 0%, bias faults require added observers, and the benchmark is released.","lead":"This paper benchmarks spacecraft fault-tolerant controllers on a strict settled-gate metric requiring sustained 0.2-degree pointing accuracy on faults absent from training. A smart generalist might read it to see why pure learning fails in safety-critical systems and how hybrid designs can succeed where both classical and end-to-end methods do not.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark splits and 6-DOF Basilisk fidelity may not capture real actuator/sensor fault generalization","rationale":"The reader's weakest_assumption directly identifies the same load-bearing external-validity risk; no internal inconsistency in the reported metrics or construction is visible from the abstract/claim, and the paper's release of the benchmark plus Wilson intervals already mitigates some reproducibility concerns.","tokens_in":1923,"tokens_out":322,"duration_ms":15281,"concrete_test":"Re-run the n=500 episode evaluation after injecting a simple time-varying gain drift (linear ramp over 30 s within the same sign/gain ranges) into the Basilisk plant; if the structured controller success rate falls below 80% while the oracle remains high, the disjoint-split generalization claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (structured estimate-then-control at 97.8%/94.4% on unseen sign/gain faults, oracle-close while unstructured methods hit 0%) depends on the claim that disjoint splits over inertia/gain/sign/bias plus the 0.2 deg settled-gate dwell in Basilisk 6-DOF are sufficient to test generalization. Real faults often include time-varying parameters, sensor-actuator coupling, unmodeled nonlinearities (e.g., friction, temperature drift), or combined fault modes absent from the disjoint construction. The abstract notes the bias wall and the need for fusion on sensor bias, but does not demonstrate that the chosen splits close the sim-to-real gap for the reported performance gap.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents an empirical benchmark for spacecraft fault-tolerant control (FTC) in a 6-DOF Basilisk simulation. It defines success via a settled-gate criterion (pointing held within 0.2 deg over a dwell window on true state) and uses train/test splits disjoint in inertia, gain, sign pattern, and bias. Across PD/PID, classical adaptive, Nussbaum-gain, end-to-end RL, and structured estimate-then-control methods (with a learned recurrent module inferring actuator gain for an analytic law), it reports 0% success for fault-unaware and unstructured methods, 55.2% for classical adaptive on gain faults, 45.2%/3.2% for Nussbaum, 97.8%/94.4% for the structured method on sign/gain faults (near oracle), 0% on constant bias for all methods, and 59.4% recovery on bias faults after adding a self-correcting disturbance observer. Sensor-fault regimes are classified similarly, emphasizing fusion needs, with the benchmark released for reproducibility (n=500 episodes per cell, Wilson intervals).","tokens_in":2075,"tokens_out":823,"duration_ms":37995,"significance":"If the benchmark protocol holds, the work supplies a standardized, statistically grounded testbed that isolates the contribution of structure versus pure learning capacity in FTC. It supplies falsifiable, held-out performance numbers and identifies a concrete failure mode (constant bias) that requires observer fusion, which could steer the field away from end-to-end RL toward hybrid designs. The release of the one-command reproduction setup and settled-gate definition is a concrete strength for cumulative progress.","major_comments":[{"comment":"Benchmark construction paragraph (and abstract): the headline performance gap (structured method at 97.8%/94.4% vs. unstructured at 0%) is interpreted as evidence for what 'actually works' on unseen faults. This interpretation rests on the claim that the chosen disjoint splits over inertia/gain/sign/bias plus the 0.2 deg settled-gate dwell are sufficient to capture relevant generalization challenges. The splits exclude time-varying parameters, sensor-actuator coupling, and combined fault modes that the skeptic note correctly flags as common in real spacecraft; without additional experiments or justification that these omissions do not alter the ranking, the external validity of the reported gap remains load-bearing and unverified.","section":"benchmark construction paragraph"},{"comment":"Disturbance-observer section (following the bias-wall result): the claim that the observer 'is self-correcting for gain-estimate error' and recovers 59.4% bias faults 'with no sign/gain regression' is central to moving the bias class off zero. The manuscript must supply the explicit coupling equations or ablation that verifies the self-correction property; absent that, the 59.4% figure cannot be assessed for robustness to the gain-inference errors already present in the structured controller.","section":"disturbance-observer section"}],"minor_comments":[{"comment":"The exact recurrent-module architecture (hidden size, input features, training objective) for the gain-inference component should be stated with a diagram or pseudocode so that the 97.8% result is one-command reproducible.","section":null},{"comment":"Table or figure reporting the per-cell Wilson intervals should include the raw success counts alongside the percentages to allow readers to recompute the intervals.","section":null},{"comment":"The 0.2 deg dwell criterion is introduced without reference to typical spacecraft pointing budgets; a short justification or sensitivity check would strengthen the settled-gate definition.","section":null}],"recommendation":"major_revision","confidential_remarks":"The title 'What Actually Works' risks overstating sim-to-real transfer; the manuscript is careful in the abstract but the title could be softened. Citation pattern is appropriate for an empirical benchmark paper."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the thoughtful and constructive report. We address each major comment below with clarifications on the benchmark design and the observer implementation. Where the manuscript requires additional detail or discussion to strengthen external validity claims and verify the self-correction property, we indicate the planned revisions.","responses":[{"response":"The benchmark deliberately isolates generalization to unseen but constant fault parameters (inertia, gain, sign pattern, bias) under a settled-gate criterion evaluated on true state, precisely to separate the contribution of controller structure from pure learning capacity. The reported gap demonstrates that unstructured methods fail under these controlled conditions while the structured estimate-then-control approach succeeds. We do not claim the ranking is invariant to time-varying faults, sensor-actuator coupling, or combined modes; those regimes lie outside the current scope. In revision we will expand the discussion section to explicitly state these scope limitations, reference the skeptic note, and outline how the released one-command reproduction setup can be extended to test the omitted regimes. No new experiments are added, as the core contribution is the controlled isolation already performed.","revision_made":"partial","referee_comment":"[benchmark construction paragraph] Benchmark construction paragraph (and abstract): the headline performance gap (structured method at 97.8%/94.4% vs. unstructured at 0%) is interpreted as evidence for what 'actually works' on unseen faults. This interpretation rests on the claim that the chosen disjoint splits over inertia/gain/sign/bias plus the 0.2 deg settled-gate dwell are sufficient to capture relevant generalization challenges. The splits exclude time-varying parameters, sensor-actuator coupling, and combined fault modes that the skeptic note correctly flags as common in real spacecraft; without additional experiments or justification that these omissions do not alter the ranking, the external validity of the reported gap remains load-bearing and unverified."},{"response":"The disturbance observer estimates the net residual (including any mismatch from the recurrent gain inference) and feeds a corrective term into the analytic control law, making the overall loop robust to gain-estimate error by construction. In the revised manuscript we will insert the explicit coupling equations between the gain-inference module output and the observer dynamics, together with an ablation that reports success rates when the observer is disabled versus enabled under the same gain-inference errors. This will allow direct assessment of the self-correction claim and the 59.4% figure.","revision_made":"yes","referee_comment":"[disturbance-observer section] Disturbance-observer section (following the bias-wall result): the claim that the observer 'is self-correcting for gain-estimate error' and recovers 59.4% bias faults 'with no sign/gain regression' is central to moving the bias class off zero. The manuscript must supply the explicit coupling equations or ablation that verifies the self-correction property; absent that, the 59.4% figure cannot be assessed for robustness to the gain-inference errors already present in the structured controller."}],"tokens_in":1799,"tokens_out":634,"duration_ms":27353,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing here is that the settled-gate benchmark with disjoint splits across inertia, gain, sign pattern, and bias produces clear separation: the hybrid recurrent-estimate plus analytic law hits 97.8% and 94.4% on held-out sign and gain faults while end-to-end RL and PD/PID sit at zero. Classical adaptive does better on sign but only 55% on gain.\n\nThe paper does the empirical side cleanly. It defines success as sustained 0.2-degree pointing over a dwell window on the true state, runs 500 episodes per cell with Wilson intervals, and tests a range of controllers including Nussbaum-gain. It flags the constant-bias wall for every integral-free method and shows the disturbance observer recovers 59% on bias without regressing the other cases. Releasing the Basilisk setup is the right move for a benchmark.\n\nThe soft spot is the generalization claim. The disjoint parameter splits test one kind of out-of-distribution behavior, but real spacecraft faults often include time variation, sensor-actuator coupling, friction, or combined modes that the current construction does not stress. The abstract already notes the bias limitation, so the work is not overselling, but the sim-to-real distance still needs more discussion.\n\nThis is for people in spacecraft attitude control who want concrete numbers on hybrid versus unstructured designs rather than another simulation win on narrow faults. The thinking is straightforward and the failures are reported, so it deserves referee time even if the numbers get pushed on in review.","headline":"A solid empirical benchmark showing structured hybrid controllers handle unseen sign/gain faults far better than end-to-end RL or basic adaptive laws, with honest reporting on the bias failure mode.","tokens_in":2575,"tokens_out":388,"would_cite":false,"duration_ms":19694,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A hybrid controller that learns actuator gain online and feeds an analytic law recovers spacecraft pointing on unseen sign and gain faults where pure learning and classical methods score zero.","keywords":["fault-tolerant control","spacecraft attitude control","adaptive control","reinforcement learning","benchmark","actuator faults","settled gate","disturbance observer"],"falsifier":"Re-running the full set of controllers on physical hardware or on a simulation whose dynamics or fault statistics differ from the Basilisk splits would show whether the reported performance ordering persists.","tokens_in":2803,"feed_emoji":"🚀","tokens_out":799,"duration_ms":18373,"temperature":0.7,"pith_summary":"The paper constructs a benchmark that demands sustained pointing accuracy within 0.2 degrees over a dwell window on actuator faults absent from training data. It evaluates controllers across disjoint splits in inertia, gain, sign pattern, and bias using a 6-DOF Basilisk simulation and Wilson intervals. Pure end-to-end learning and fault-unaware PD/PID controllers achieve zero success under this settled-gate metric. Classical adaptive laws handle sign faults but reach only 55 percent on gain faults, while a structured design pairing a learned recurrent gain estimator with an analytic law reaches 97.8 percent and 94.4 percent on those classes. A disturbance observer added to the same structure recovers 59.4 percent of constant bias faults without degrading prior performance.","feed_headline":"Hybrid controller recovers 94 percent on unseen spacecraft gain faults","feed_subtitle":"Learned gain estimator feeding analytic law succeeds where pure learning and classical adaptive methods score zero under settled-gate test.","key_machinery":"The estimate-then-control structure in which a learned recurrent module infers actuator gain online and supplies it to an analytic control law, extended by a disturbance observer for bias.","core_discovery":"The central claim is that an estimate-then-control architecture, in which a recurrent module infers actuator gain from state history and supplies the estimate to a model-based law, achieves high success on held-out sign and gain faults under a settled-gate criterion, while unstructured learned and classical controllers remain at zero success and constant bias remains unsolved without an additional observer.","pith_inferences":["The same estimate-then-control pattern may transfer to other parametric-uncertainty problems where an analytic law is available but parameters must be inferred online.","The zero success on bias faults indicates that integral action or equivalent observers are required for constant disturbances in any fault-tolerant spacecraft controller.","Releasing the benchmark with its settled-gate definition and disjoint splits allows direct comparison of future methods on the same generalization test.","The requirement for sensor fusion on bias cases suggests that fault detection in attitude control will often need multiple independent measurements rather than observers alone."],"forward_implications":["Fault-unaware PD/PID and from-scratch end-to-end reinforcement learning achieve zero success on the settled gate for any tested fault class.","Classical adaptive laws resolve sign faults but reach only 55.2 percent on gain faults and 45.2 percent or lower with Nussbaum-gain variants.","Constant additive bias defeats every controller including the privileged gain oracle because an integral-free law cannot reject a persistent disturbance.","A disturbance observer recovers 59.4 percent of held-out bias faults when composed with the gain estimator and produces no regression on sign or gain cases.","Sensor bias is unobservable from corrupted measurements alone and therefore requires fusion rather than a standalone observer."],"fun_headline_variants":["Structured controller with recurrent gain estimator recovers 94 percent on faults","Recurrent inference of actuator gain feeds analytic law for 94 percent fault recovery","Estimate then control with learned gain module succeeds on unseen spacecraft faults","Learned recurrent gain estimator plus analytic law achieves 94 percent on gain faults","Hybrid estimate control design recovers pointing under held out gain faults at 94 percent"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 6-DOF Basilisk simulation with the chosen disjoint splits in inertia, gain, sign pattern, and bias plus the 0.2 degree settled-gate dwell criterion captures the generalization challenges of real spacecraft actuator and sensor faults.","fun_headline_variants_meta":{"raw":{"variants":["Structured controller with recurrent gain estimator recovers 94 percent on faults","Recurrent inference of actuator gain feeds analytic law for 94 percent fault recovery","Estimate then control with learned gain module succeeds on unseen spacecraft faults","Learned recurrent gain estimator plus analytic law achieves 94 percent on gain faults","Hybrid estimate control design recovers pointing under held out gain faults at 94 percent"]},"model":"grok-4.3","cost_usd":0.005522,"raw_usage":{"total_tokens":2726,"prompt_tokens":819,"num_sources_used":0,"completion_tokens":93,"cost_in_usd_ticks":55224500,"prompt_tokens_details":{"text_tokens":819,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1814,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":819,"tokens_out":93,"duration_ms":14363,"temperature":1.0,"reasoning_tokens":1814,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-25T21:15:37.067613+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-running the full set of controllers on physical hardware or on a simulation whose dynamics or fault statistics differ from the Basilisk splits would show whether the reported performance ordering persists.","supporting_citations":[],"review_version":1}