{"id":"adddc0d2-6f75-484d-af4f-dc408e6f0923","arxiv_id":"1908.01876","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A stability basin computed from subject-specific pendulum models and data-driven control-input bounds predicted 90.9% of step failures, 92.2% of sit failures, and 95.7% of successful sit-to-stand trials in an 11-person cable-pull experiment.","lead":"Researchers tested a model-based 'stability basin' that defines when a person doing a sit-to-stand can stay stable under a cable pull. In 11 healthy adults it classified about 9 in 10 failed trials and about 19 in 20 successful trials, beating three comparison methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Failure accuracy may be inflated by late failures: since natural/MT step onsets average >100% STS and t_f is reset to 1, an exit from the terminal set XT at t=1 can count as a prediction, so the reported >90% is not established.","rationale":"The reader's conditional verdict is appropriate, and several of the reader's listed concerns are real. I do not see a reason to change the category of the verdict, but I would sharpen the most load-bearing issue. The paper's strongest independent evidence is the failure-trial prediction rate, yet the evaluation as described may count terminal-set exclusion at t=1 as a 'prediction' for the many failures whose onset occurs after the nominal trial end. The paper currently gives no first-exit-time distribution or lead-time analysis, so the headline 'over 90% failure prediction' cannot be interpreted as demonstrating that the SB boundary is crossed during motion in advance of the failure event. This is distinct from the reader's weakest assumption about input-envelope and standing-set coverage: even if the BFF+FB envelope and XT are perfectly calibrated, the reported accuracy could still be inflated by the timing convention. The proposed reanalysis of the existing public data would settle the question directly. If the lead times are substantial and accuracy survives the exclusion of late-failure trials, then the central claim is credible; if not, the claim should be downgraded. The paper's open code and data are a genuine strength and make this check straightforward.","tokens_in":17292,"tokens_out":14970,"duration_ms":177997,"concrete_test":"Using the public GitHub data, compute for every failure trial the first time step at which the observed trajectory exits the SB, and stratify by t_f: (i) t_f < 0.9, (ii) 0.9 <= t_f <= 1.0, (iii) t_f > 1.0. Then recompute the Table 2 step/sit accuracy after excluding group (iii) and after requiring the first exit to occur at least 5% STS before t_f (with t_f capped at 1). If the pooled failure accuracy drops below roughly 90% or the median lead time between first exit and failure onset is near zero, the headline overstates the predictive content of the SB.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central validation claim rests on the 90.91% step and 92.21% sit prediction rates in Table 2. These rates count any exit from the SB before t_f as a correct failure prediction, and for trials whose defined failure onset falls after the normalized trial end, t_f is reset to 1 and the trajectory is checked all the way to t=1. Table 1 reports mean step onsets of 107.27% (natural), 114.30% (MT), and 95.04% (QS) of normalized trial time, and mean sit onsets of 96.16%, 99.38%, and 85.26%. Thus for natural and MT trials in particular, a large share of 'failure predictions' may reduce to the observation that the state at t=1 is not in XT, the terminal standing set built from successful trials' final states. Because XT is by construction the set of successful final states, a perturbed trial that has not reached standing by the nominal end is almost automatically outside the SB at t=1. The paper does not report the distribution of first-exit times or the lead time between the first SB exit and t_f, so it is impossible to tell whether the SB boundary crossing precedes the failure by a meaningful margin or merely coincides with the terminal-set check. This is more load-bearing than the envelope-coverage concern because it affects the interpretation of the only independent validation signal (failure trials) rather than a parametric assumption that could be explored by sensitivity analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a data-driven method to compute subject- and strategy-specific Stability Basins (SBs) for the Sit-to-Stand (STS) task, using a Telescoping Inverted Pendulum Model, bounded feedforward-plus-feedback (BFF+FB) control bounds estimated from successful trials (Eqs. 6-8), and a standing zonotope XT built from successful final states. The SB is computed by backwards reachability with CORA. The method is validated on an 11-subject perturbative STS experiment with 948 trials, including 198 cable-pull-induced failures (steps and sits). The authors report that SBs predict 90.91% of steps and 92.21% of sits before failure onset, and 95.73% of successful trials in a leave-one-out procedure, outperforming LQR, FF+FB, and naive trajectory-envelope baselines. The central claim is that a stability basin constructed only from successful-trial kinematics and control bounds can classify whether a perturbed STS trial will end in a step or sit before the failure occurs.","tokens_in":17595,"tokens_out":6582,"duration_ms":74127,"significance":"If the validation is sound, this is a significant methodological contribution: it provides the first experimental validation of a reachability-based Stability Basin for an aperiodic, non-periodic movement, and it does so with a computationally tractable pipeline (about 0.76 s per SB) that uses only successful trials, which is clinically attractive. The comparison against LQR, FF+FB, and a naive envelope is a useful benchmark, and the data/code release (GitHub) supports reproducibility. However, the central validation claim is currently not established at the level the abstract implies: the pooled 'over 90%' failure rate is not achieved in the Momentum-Transfer condition, and a substantial fraction of failure trials have their defined onset after the normalized trial end, where the terminal-set check may make failure predictions nearly tautological. These issues are addressable with additional analysis of the existing data, so the contribution is potentially strong but needs revision.","major_comments":[{"comment":"The resetting of tf to 1 for failures whose onset occurs after the trial end, combined with a standing zonotope XT built exclusively from successful final states, makes the terminal check x(1)∈XT nearly tautological for late-onset failures. Table 1 reports mean step onsets of 107.27% (natural) and 114.30% (MT) of normalized trial time with standard deviations of 20.92 and 32.66; assuming approximate normality, roughly 60-70% of natural and MT step trials have onset after t=1. For those trials, the prediction is based on the full trajectory up to t=1, and any exit from the SB—including an exit at t=1 because the subject has not reached the successful standing set—counts as a correct failure prediction. Because the paper does not report first-exit times or the lead time between the first SB exit and tf, the reported 90.91% step and 92.21% sit rates do not establish that the SB boundary is crossed before the behavioral failure. Please report the fraction of failure trials with tf>1, the distribution of first-exit times, and re-compute accuracy using only trials with tf≤1 and with a minimum lead-time requirement.","section":"Sec. 2.4.1, Table 1"},{"comment":"The headline claim of 'over 90% accuracy' for failure predictions is a pooled rate that is not achieved in the Momentum-Transfer condition. Table 2 reports 33/40 (82.5%) steps and 17/19 (84.75%) sits for MT, while natural and QS strategies are at least 92%. The abstract and introduction state an unqualified 'over 90%' accuracy, which is unsupported for a full third of the tested strategy conditions. Please report per-strategy rates with confidence intervals and qualify the abstract claim accordingly.","section":"Abstract, Table 2"},{"comment":"The input envelope used to construct the SB is fit only to the non-perturbed portions of successful cable-pull trials; the reactive control inputs during the 250 ms perturbation are discarded in the inverse-dynamics computation. Thus the BFF+FB set does not actually cover the control inputs a human uses to recover from the perturbation, which is the central modeling assumption the validation is intended to test. If the envelope is too narrow during the perturbation window, the SB will be overconservative and failure predictions will be inflated. Please report how many successful trials contribute to the envelope at each time step, and compare failure-prediction rates for SB variants that include perturbed-window inputs from successful CP trials or use an explicitly enlarged input bound during perturbation.","section":"Sec. 2.3.1, Remark 3"},{"comment":"All prediction rates are pooled across 11 subjects and across trials without accounting for within-subject correlation or reporting subject-level variability. Because each subject contributes many trials (for example, 591 cable-pull trials across subjects), a small number of subjects could dominate the aggregate. Please provide per-subject accuracy tables, confidence intervals (e.g., by bootstrap or a mixed-effects model), and the range of per-subject success and failure prediction rates.","section":"Sec. 2.4.3, Table 2"}],"minor_comments":[{"comment":"The leave-one-out description is ambiguous: 'After forming the standing set, we leave one successful trial out...'—it is not clear whether the left-out trial's final state is included in the standing zonotope XT. If it is, the t=1 membership of that trial is guaranteed by construction; please clarify and, if appropriate, also leave the trial out of the XT construction.","section":"Sec. 2.4.3"},{"comment":"The LQR weighting matrices Q=I and R=10^-4 I are reported as 'found empirically to produce the best results,' but the data used for this tuning are not specified. If the weights were tuned on the validation set, the comparison in Table 4 is biased; please state the tuning procedure and whether the validation data were used.","section":"Sec. 2.5.1"},{"comment":"The statement that the SB 'estimates the stable region with over 45% more accuracy' is not defined precisely. The table mixes success and failure rates with different denominators; please report a single pre-specified accuracy metric (e.g., balanced accuracy or Matthews correlation coefficient) and give per-strategy values.","section":"Table 4, Sec. 3"},{"comment":"The paper would benefit from confidence intervals or Bayesian credible intervals for the main prediction rates; with only 11 subjects and modest numbers of failures per strategy, the reported percentages (e.g., 17/19 sits for MT) have wide uncertainty.","section":"General"},{"comment":"The phrase 'where the times at which the peak horizontal COM velocities occur are coincident' should specify whether the alignment is done on filtered data and whether the same alignment is used for all strategies; the QS procedure uses vertical COM velocity, which is a different criterion and could affect comparability.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about late failures is well founded: Table 1 shows mean step onsets above 100% STS for natural and MT trials, so a large share of 'failure predictions' may be driven by the terminal check against a set built from successful final states. The paper's main contribution is potentially important, but the validation claims need reanalysis with first-exit-time distributions and per-strategy reporting before publication. I would also verify that the GitHub repository contains the exact analysis code and data needed to reproduce Tables 1-4."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe bottom line: this is the first real perturbative validation of the Stability Basin idea, and the BFF+FB controller construction is a genuine step forward. But the headline 90% failure-prediction accuracy is not established, because for natural and momentum-transfer strategies a large fraction of step failures occur after the normalized trial end, and the method then counts an exit from the terminal standing set at t=1 as a prediction. That is close to a tautology.\n\nWhat the paper does well: it builds on Shia et al. but replaces the single LQR controller with a data-driven envelope of feedforward plus linear feedback inputs, which is a sensible way to avoid overfitting a single controller to human data. The experiment is carefully done: 11 subjects, three strategies, cable pulls with graded forces, 198 observed failures. The failure trials are genuinely independent of the basin construction, and the leave-one-out scheme for successful trials is sound. The comparison to LQR, FF+FB, and a trajectory-envelope method shows the proposed basins are far less conservative, which is a real improvement.\n\nThe soft spot: the evaluation protocol for failures lets t_f=1 when the step or sit occurs after the segmented trial end. For natural and MT steps the onset averages 107% and 114% of normalized trial time. Since the SB at t=1 is essentially the standing zonotope XT built from successful final states, any perturbed trial that hasn't reached standing by the nominal end will be flagged as a failure, regardless of whether the state is on an unstable trajectory. The paper does not report first-exit times or the lead time between exit and onset, so we can't tell how many of the 'correct' predictions are just terminal-set checks. This could easily account for the near-perfect natural-step accuracy (97.6%), and it weakens the central independent-validation claim. A reanalysis that excludes late failures or reports exit-time distributions is needed.\n\nMinor points: the left-out successful trial's final state is still in XT, which slightly inflates success accuracy; the 5% zonotope expansion and the step criterion are post hoc; and the GitHub link has no commit hash. None of these are fatal.\n\nVerdict: the BFF+FB method is worth building on, and the experiment is worth publishing, but the current validation does not support the advertised accuracy. A serious referee should see this under major revision, with the late-failure analysis front and center. I'd bring the paper to a reading group to talk about how easy it is to fool yourself with terminal-set validation.","headline":"A promising first validation of the Stability Basin idea, but the 90% failure-accuracy claim is undermined by a terminal-set loophole for late failures.","tokens_in":18198,"tokens_out":4896,"would_cite":false,"duration_ms":50825,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 'stability basin' built from a person's own motion predicts sit-to-stand failure before it happens.","keywords":["stability basin","sit-to-stand","fall risk","reachability analysis","bounded feedforward-feedback control","perturbative balance experiment","center-of-mass dynamics"],"falsifier":"Run cable-pull trials that deliberately push a subject through states the basin labels unsafe while coaching them to recover with a strategy that never appears in the successful trials used to build the basin (for example, a quick corrective step); if those trials are recovered without stepping or sitting, the envelope under-covers the true stable region. A cheaper check is to rebuild each basin with the standing set expanded by 0% and by 20% instead of 5% and see whether the reported accuracy survives; if it changes sharply, the headline numbers depend on that tuning choice.","tokens_in":17017,"feed_emoji":"🪑","tokens_out":12793,"duration_ms":118410,"temperature":0.7,"pith_summary":"This paper tries to establish that a person's resilience during sit-to-stand can be predicted ahead of time by a computed 'stability basin' — the set of body states from which, under their own control strategy, they can reach standing without stepping or sitting back down. The study built individualized basins from the motion of 11 people and checked them against an experiment that pulled subjects forward or backward by a cable while they rose. The basins flagged over 90% of actual failures before the step or sit occurred and correctly accepted over 95% of successful trials. Such a result matters because current clinical fall-risk tests have weak predictive power, whereas an accurate basin could give therapists a quantitative, subject-specific measure of stability and guide assistive devices.","feed_headline":"Stability basin predicts sit-to-stand failures with over 90% accuracy","feed_subtitle":"A model built from a person's own successful sit-to-stand trials predicts when a cable pull forces a step or sit.","key_machinery":"The central object is the Stability Basin, defined for each subject and sit-to-stand strategy as the set of times and body states from which the model can reach standing without violating control limits. Body motion is represented by a Telescoping Inverted Pendulum Model: a point mass at the center of mass with separate horizontal and vertical position and velocity states. Two data-driven objects make the basin match the individual: a 'bounded feedforward plus feedback' (BFF+FB) controller, whose feedforward term may range over an envelope fitted by quadratic programming to the inverse-dynamics inputs of all successful trials, and a standing set formed as a zonotope — a symmetric polytope described by a center and generators — that encloses the final states of successful trials, with each generator expanded by 5%. Reachability analysis flows the standing set backward through time under the pendulum dynamics and input envelope, yielding time-indexed basins; a trial is predicted to fail if its observed trajectory exits the basin before the step or sit event, and to succeed if it never exits.","core_discovery":"The paper's central claim is that a Stability Basin — the set of body states from which a person can finish standing under their own control strategy — can be computed from successful trials alone and then, in a perturbed trial, correctly separates recoverable perturbations from those that force a strategy switch. In the experiment, cable pulls applied to the waist of 11 subjects during natural, momentum-transfer, and quasi-static sit-to-stand motions produced 198 trials in which the subject stepped or sat back down; such events define failure. State trajectories from motion capture were checked against each subject's basin, and the basin predicted 110 of 121 steps and 71 of 77 sits before their onset (failure accuracy 90.9% and 92.2%, combined 91.4%), while 718 of 750 successful trials stayed inside the basin (95.7%). Three alternative constructions — a linear-quadratic regulator controller, a single feedforward-plus-feedback controller, and a naive envelope around successful trajectories — predicted successful trials only 9.5%, 46.8%, and 16.7% of the time, respectively, meaning they grossly under-estimate the stable region. The paper concludes that the bounded, data-driven input model is what makes the basin both accurate and not overly conservative.","pith_inferences":["The paper's 17 misclassified failures lie, on average, closer to successful trajectories and occur earlier in the pull than correct failure predictions, which suggests the basin edge is a gray zone rather than a hard boundary — a possibility the paper does not test.","Varying the 5% expansion of the standing set (for instance, 0% and 20%) would show how much of the reported accuracy depends on that tuning constant; the paper does not report such a sensitivity check.","If a person can recover with a strategy never shown in the successful trials used to fit the input envelope — such as a deliberate lunge — the method would label it a failure; a protocol that elicits and tests such alternate recoveries could expose under-coverage of the envelope.","The same backwards-reachability pipeline should transfer to other aperiodic tasks, such as step-ups or reaching, provided the completion set and failure definitions are equally well specified; this is a natural next application not addressed in the paper."],"forward_implications":["The method classifies whether a perturbed sit-to-stand trial will end in a step or sit before that event occurs, using only center-of-mass kinematics from the trial.","Because the basin is built exclusively from successful trials, it can be applied to frail or fall-prone individuals without ever perturbing them to the point of falling.","Each strategy-specific basin computes in under a second on a laptop, which makes the approach fast enough for clinical screening or for near-real-time feedback.","Repeated assessments could track whether an individual's basin shrinks or grows over time, giving a quantitative, task-specific fall-risk monitoring tool.","Wearable assistive devices could use the basin as a safety constraint or optimization target, adjusting support when the user is about to exit the stable region."],"supporting_citations":[{"why":"It introduced the Stability Basin concept and the LQR baseline against which the proposed method is compared.","marker":"[31]"},{"why":"It supplied the Telescoping Inverted Pendulum Model used to represent each subject's center-of-mass motion.","marker":"[35]"},{"why":"It provides the anthropometric segment parameters used to estimate center-of-mass trajectories from the marker data.","marker":"[34]"},{"why":"It supplies the inverse-dynamics equations used to compute the control inputs of successful trials for fitting the input envelope.","marker":"[38]"},{"why":"It provides the zonotope construction used to build the standing set from the final states of successful trials.","marker":"[40]"},{"why":"It supplies the reachability-analysis software used to compute the backwards-reachable stability basins.","marker":"[41]"},{"why":"It defines stepping and sitting as the failure modes that mark a switch in control strategy during sit-to-stand.","marker":"[30]"},{"why":"It motivates the feedforward-plus-feedback decomposition that the bounded controller generalizes.","marker":"[36]"}],"fun_headline_variants":["Stability basin predicts fall-triggering perturbations in sit-to-stand","Cable-pull experiments validate stability basin for fall risk","Model-based basin identifies when sit-to-stand fails","Stability basin beats three alternatives in predicting falls"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole prediction rests on the assumption that the forces a person applied and the standing poses observed during successful trials cover every recovery strategy that person could actually use; if a person recovers using an action outside that observed range, the computed basin will flag them as unstable even though they are not.","fun_headline_variants_meta":{"raw":{"variants":["Stability basin predicts fall-triggering perturbations in sit-to-stand","Cable-pull experiments validate stability basin for fall risk","Model-based basin identifies when sit-to-stand fails","Stability basin beats three alternatives in predicting falls"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1450,"prompt_tokens":1101,"completion_tokens":349,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":717,"completion_tokens_details":{"reasoning_tokens":281}},"tokens_in":717,"tokens_out":349,"duration_ms":4250,"temperature":1.0,"reasoning_tokens":281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:01:31.646556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run cable-pull trials that deliberately push a subject through states the basin labels unsafe while coaching them to recover with a strategy that never appears in the successful trials used to build the basin (for example, a quick corrective step); if those trials are recovered without stepping or sitting, the envelope under-covers the true stable region. A cheaper check is to rebuild each basin with the standing set expanded by 0% and by 20% instead of 5% and see whether the reported accuracy survives; if it changes sharply, the headline numbers depend on that tuning choice.","supporting_citations":[{"cited_title":"& Vasudevan, R","cited_arxiv_id":null,"evidence_quote":"It introduced the Stability Basin concept and the LQR baseline against which the proposed method is compared."},{"cited_title":"& Cappozzo, A","cited_arxiv_id":null,"evidence_quote":"It supplied the Telescoping Inverted Pendulum Model used to represent each subject's center-of-mass motion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the anthropometric segment parameters used to estimate center-of-mass trajectories from the marker data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the inverse-dynamics equations used to compute the control inputs of successful trials for fitting the input envelope."},{"cited_title":"& Krogh, B","cited_arxiv_id":null,"evidence_quote":"It provides the zonotope construction used to build the standing set from the final states of successful trials."},{"cited_title":"An introduction to CORA 2015","cited_arxiv_id":null,"evidence_quote":"It supplies the reachability-analysis software used to compute the backwards-reachable stability basins."},{"cited_title":"Krebs, D","cited_arxiv_id":null,"evidence_quote":"It defines stepping and sitting as the failure modes that mark a switch in control strategy during sit-to-stand."},{"cited_title":"The relative roles of feedforward and feedback in the control of rhythmic movements","cited_arxiv_id":null,"evidence_quote":"It motivates the feedforward-plus-feedback decomposition that the bounded controller generalizes."}],"review_version":1}