{"id":"f5153a40-ab5f-4d58-b21c-5fcaf469f3d2","arxiv_id":"2504.17201","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A three-mode multiple-model Kalman filter using only joint encoders simultaneously detects leg collisions and estimates external forces, enabling safer quadrupedal locomotion.","lead":"This paper combines an interacting multiple-model Kalman filter with robot dynamics models to detect leg collisions and estimate contact forces using only joint encoders. The estimates drive reflex, admittance, and model predictive controllers that reduce trip impact in quadrupedal walking.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Collision/stance discrimination rests on an unmeasured terrain prior: a vertical-force collision (e.g., striking the top of a low step) is assigned to the stance cone, so the 'joint encoders and dynamics only' claim is narrower than stated.","rationale":"The reader's weakest assumption and my concern coincide: the stance/collision distinction is enforced by fixed vertical vs horizontal noise cones, so collision forces with vertical direction are misclassified. This is the single most load-bearing issue because the central contribution is simultaneous mode identification and force estimation from proprioception alone; if mode identification depends on an external terrain prior, the headline claim must be qualified. I do not see an internal inconsistency or a reason to reject: the estimator is a reasonable IMM application, Eq. (9)-(17) are standard, the conical pseudo-measurement noise is a legitimate design choice, and the simulation benchmark (89 collisions, Table II) and hardware demonstration provide real supporting evidence for the frontal-collision regime. The concern is a scope limitation, it is explicitly acknowledged in the Conclusion, and it is addressable by adding kinematic cues such as foot velocity or by adapting the cone online. For these reasons the reader's CONDITIONAL verdict remains appropriate; no verdict change is needed.","tokens_in":9260,"tokens_out":4824,"duration_ms":49646,"concrete_test":"In the Gazebo environment, add a low horizontal step ahead of the robot such that the swing foot contacts its top surface during the downward phase of swing, producing a predominantly vertical ground-truth contact force inside the stance cone F_v. Run the IMM-MBKO estimator and record the mode probabilities and the reflex trigger. If the collision-mode probability stays below the stance-mode probability during contact and no reflex is triggered, the cone-prior concern is confirmed; if the filter still detects collision, the dynamics/GM information is doing more discrimination than the current ablation shows.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The filter cannot distinguish collision from stance using the process model: in Eq. (10), both stance (k=2) and collision (k=3) use S(k)=I, and the GM dynamics are identical. The only separation is the mode-dependent measurement noise in Eq. (15) and Fig. 1: a pseudo-force inside the vertical cone F_v is treated as stance, a force inside the horizontal cone F_h as collision. Because the cone orientation is fixed a priori, the classification is a terrain-shape prior, not a quantity measured by joint encoders. A collision that produces a predominantly vertical contact force, such as a swing foot descending onto the top surface of an unseen low step or obstacle, lies inside F_v; the stance filter will have small measurement noise, the likelihood in Eq. (17e) will favor stance, the collision probability will remain low, and the reflex/admittance/MPC responses will not trigger. This is exactly the limitation the authors acknowledge in the Conclusion ('some rough prior knowledge of terrain shape is required'), and the simulation/hardware evidence only exercises roughly horizontal obstacle impacts. The failure does not invalidate the demonstrated performance on frontal collisions, but it does narrow the abstract's claims of detection using 'joint encoder information and the robot dynamics only' and of invariance: the estimator is invariant to gait pattern, but not to terrain-shape prior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an interacting multiple-model Kalman filter (IMM-KF) for simultaneous external force estimation and contact-mode detection in quadrupedal locomotion. The estimator uses generalized momentum dynamics, a mode-dependent process model, pseudo-force measurements derived from the robot dynamics, and mode-dependent measurement noise based on friction-cone orientation. Three contact modes are considered: swing, stance, and collision. The estimated mode probabilities and forces drive a reflex step-height increase, a swing-leg admittance controller, and a force-adaptive model predictive controller. Validation is performed in Gazebo against three baselines across 89 collisions and on a Unitree A1 robot in a single obstacle scenario.","tokens_in":9504,"tokens_out":4726,"duration_ms":42991,"significance":"If the claims hold, the proposed method offers a proprioceptive, gait-invariant alternative to vision-based collision detection for quadrupedal robots, with the practical advantage of not requiring force sensors or terrain perception at runtime. The paper contributes a principled multi-mode filtering architecture that extends momentum-based observers to simultaneous detection and force estimation, and it demonstrates the integration of these estimates into a reflex controller and an MPC. The hardware demonstration and the comparison against three baselines are valuable. However, the headline claim of using 'joint encoder information and the robot dynamics only' is partially undercut by the terrain-shape prior embedded in the friction-cone measurement model, and the comparative evaluation relies on single-run point estimates without statistical testing. With appropriate caveats and a strengthened evaluation, the method could be a useful addition to the legged-locomotion literature.","major_comments":[{"comment":"The discrimination between stance and collision relies entirely on the pre-defined orientation of the friction cones F_v and F_h. While the process model in Eq. (9)-(10) distinguishes swing from the other modes, it does not distinguish stance from collision because S(2)=S(3)=I. Thus the only separation between stance and collision is the prior assumption that collision forces are horizontal and stance forces are vertical. A collision force with a large vertical component, such as a swing foot striking the top of a low step, falls inside F_v, the stance filter is favored in the likelihood update of Eq. (17e), and the reflex, admittance, and MPC responses will not trigger. This directly contradicts the abstract's claim that the method uses 'joint encoder information and the robot dynamics only' and narrows the invariance claim to a fixed terrain-shape prior. The paper acknowledges this in the Conclusion, but the abstract and the architecture description do not carry the caveat, and the simulations and hardware experiments exercise only roughly horizontal impacts. This limitation should be stated prominently and the claims should be adjusted accordingly.","section":"Section III-B, Eq. (15) and Fig. 1"},{"comment":"The benchmark reports one set of 89 collisions with no indication of the number of independent simulation runs, no variance or confidence intervals, and no statistical test. The difference between IMM-MBKO (85/89) and MBKO (84/89) is a single detection and is not shown to be meaningful. The same issue affects Table III, where all metrics (61 vs 71 collisions, 0.052 vs 0.036 s duration, 2.134 vs 1.15 Ns impulse) are single-run point estimates without any measure of spread. Without repeated trials or a statistical test, the claimed superiority over the baselines is not established.","section":"Section V-B, Table II"},{"comment":"The detection criterion underlying the numbers in Table II is not defined. The paper does not state the threshold on the mode probability µ(k) or the required persistence time for a collision to be counted as detected, nor how false positives and false negatives are exactly scored relative to the ground-truth contact state. This makes the benchmark results irreproducible and the comparison with the baselines difficult to interpret, especially because the baselines use a similar contact-mode logic that is described only qualitatively.","section":"Section V-B"}],"minor_comments":[{"comment":"There are minor language issues: 'ablatation studies' should be 'ablation studies', and 'independent on external sensors' should be 'independent of external sensors'.","section":"Abstract and Section I"},{"comment":"The Swing RMSE value for PM-MBKO is printed as '.56'; it should be '0.56' for consistency and clarity.","section":"Table II"},{"comment":"The symbols \\hat F^* and F^* in Eq. (19) are not defined. Please clarify what the asterisk denotes and whether the norm is computed per sample or averaged over the collision interval.","section":"Eq. (19)"},{"comment":"The cone geometry and the numerical threshold separating 'large' and 'small' measurement noise in Eq. (15) are not specified. The values v_f = 0.001/200 in Table I are ambiguous; please state how the two values map to the inside/outside of the cone.","section":"Fig. 1 and Eq. (15)"},{"comment":"The transition probability matrix sets π_2 and π_3 such that direct transitions between stance and collision have probability zero. This assumption is not justified, and it may be violated in scenarios where the robot is already in stance when a collision occurs; please comment on the sensitivity to this choice.","section":"Section V-A"}],"recommendation":"major_revision","confidential_remarks":"The paper is closely related to the authors' prior work [26], and the incremental contribution would be clearer with a direct comparison against that method rather than only against the MBKO variants. The abstract's 'joint encoder information and the robot dynamics only' claim is too strong given the terrain-shape prior in the measurement model; this should be moderated in revision. The single-run evaluation in Table II and III should be strengthened before acceptance, as the current evidence does not robustly support the performance claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read it. The core idea is a sensible three-mode IMM-KF (swing/stance/collision) that fuses the generalized-momentum process model with a pseudo-force measurement, and then feeds the mode probabilities and force estimate into a reflex step, an admittance law, and a force-adaptive MPC. That combined package is not in the earlier work from the same group [26] or elsewhere, so it earns the novelty claim. The simulation benchmark against three baselines is clear, the metrics are reasonable, and the hardware demo, while not quantitative, at least shows the reflex and balancing controllers working on a real A1. The math is internally consistent; I did not see a circular step.\n\nThe main soft spot is the one the authors half-admit in the conclusion: 'some rough prior knowledge of terrain shape is required to define the mode-dependent measurement models.' That is doing real work. The process dynamics are identical for stance and collision (both S=I), so the only thing separating them is the measurement noise model, which assumes stance forces are vertical and collision forces are horizontal. A swing foot that strikes the top of a low step or obstacle produces a predominantly vertical contact force. That lies inside the stance cone, the likelihood favors stance, the collision probability stays low, and the reflex/admittance/MPC reactions do not trigger. The demonstrated cases are frontal, roughly horizontal impacts, so the claim of detection using only joint encoders and dynamics is narrower than the abstract suggests. This is a credible limitation, not a fatal flaw, but it should be stated on page one, not just in the conclusion.\n\nTwo smaller issues: Table II reports a single run of 89 collisions with no variance or statistical test, so the performance deltas over baselines may not be as clean as they look. And the simulation uses one obstacle layout; the gait-invariance claim is untested beyond a trot. These are easily fixable in revision.\n\nBottom line: worthwhile work, honestly presented, with one load-bearing caveat that needs to be moved from the conclusion to the front and addressed in the experiments. I would send it to review; a good referee will ask for new experiments with vertical impacts and repeated trials, but the core mechanism appears sound.","headline":"Useful IMM-KF extension for proprioceptive collision detection on quadrupeds, but the stance/collision cone geometry silently imports a terrain prior that narrows the 'dynamics only' claim.","tokens_in":10102,"tokens_out":2206,"would_cite":false,"duration_ms":19690,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A mode-switching Kalman filter on joint-encoder data can simultaneously detect a leg collision and estimate the external contact force, and feeding those signals into a reflex, an admittance controller, and a force-adaptive MPC improves…","keywords":["quadrupedal locomotion","collision detection","contact force estimation","interacting multiple-model Kalman filter","momentum-based observer","proprioceptive sensing","admittance control","model predictive control"],"falsifier":"A concrete test is to command the swing foot to strike the top of a low step while the foot is still moving downward, so the contact force is mostly vertical. If the IMM-KF assigns that event high stance probability rather than collision probability and no reflex triggers, the central detection claim fails for that case; this can be checked with a force-sensor-instrumented foot or a simulation with ground-truth contact labels.","tokens_in":8996,"feed_emoji":"🦿","tokens_out":9336,"duration_ms":81385,"temperature":0.7,"pith_summary":"This paper tries to establish that a quadruped can detect an unexpected leg collision and measure the impact force at the same time, using only its joint encoders and a model of its own dynamics. The estimator is an interacting multiple-model Kalman filter: one Kalman filter per contact mode, swing, stance, and collision, mixed by a Markov chain, so the result is a probability for each mode plus a combined external-force estimate. The force entering the measurement update is a pseudo-measurement computed from the dynamics, not a sensor reading, and the modes are told apart by whether that pseudo-force falls inside a vertical or horizontal friction cone. On top of this, the paper builds a reflex that lifts the colliding foot, an admittance controller that lets the leg yield to the force, and a force-adaptive model-predictive controller for balance. If the claim holds, legged robots get gait-pattern-invariant collision handling without force sensors or terrain perception.","feed_headline":"Joint encoders alone can reveal leg collisions and impact force","feed_subtitle":"A mode-switching Kalman filter separates swing, stance, and collision, then triggers a reflex and rebalancing.","key_machinery":"The central object is the interacting multiple-model Kalman filter (IMM-KF), a bank of Kalman filters, each matched to one contact mode, whose estimates are mixed through a Markov transition matrix and whose mode probabilities are updated from the innovation likelihoods. The load-bearing mechanism inside it is the mode-dependent measurement model: a pseudo-measurement of the external force computed from the robot dynamics and contact constraint, combined with friction-cone-shaped measurement noise that gives each mode a region of trust. That cone geometry, not a force threshold or gait timing, is what lets the filter distinguish an unexpected collision from an intentional stance. The filter outputs the blended force estimate $\\hat{\\mathbf{x}}_{t|t}$ and mode probabilities $\\mu_t^{(k)}$ that the reflex, admittance, and MPC modules consume.","core_discovery":"The paper's central claim is that simultaneous collision detection and external-force estimation can be solved by a single mode-conditioned estimator, an interacting multiple-model Kalman filter (IMM-KF), using only joint encoders and robot dynamics. Three Kalman filters run in parallel, one per contact mode (swing, stance, collision), sharing generalized-momentum dynamics but differing in the selection matrix $\\mathbf{S}^{(k)}$ that connects external force to momentum, and in the measurement noise model. In the measurement update, a pseudo-force $\\mathbf{F}_{\\mathrm{pse},j} = -(\\mathbf{J}_{c_j}\\mathbf{M}^{-1}\\mathbf{J}_{c_j}^T)^{\\dagger}(\\mathbf{J}_{c_j}\\mathbf{M}^{-1}\\boldsymbol{\\tau} + \\dot{\\mathbf{J}}_{c_j}\\dot{\\mathbf{q}})$ is computed from encoder data and the contact constraint; the filter trusts that pseudo-force as a measurement of the external force only when it lies inside the mode's friction cone (vertical for stance, horizontal for collision), otherwise it treats it as large noise. The IMM combination outputs a probability for each mode and a blended force estimate. The paper further demonstrates that these two outputs can drive a reflex that lifts the foot, an admittance controller that lets the leg yield, and a model-predictive controller that treats the estimated force as a disturbance, improving impulse and balance in simulation and in a hardware collision with a table leg.","pith_inferences":["Beyond the paper, the cone-based discrimination suggests that adding contact submodes, such as slip, rear-end collisions, or two-foot impacts, may be as simple as adding new cones and relaxing the transition matrix's zero entries for direct stance-to-collision and collision-to-stance changes.","Beyond the paper, using foot velocity in the transition probabilities, a direction the authors flag for future work, could cut detection delays because a downward-moving foot striking an obstacle is kinematically different from a planted stance foot.","Beyond the paper, the pseudo-wrench expression includes moment terms, so the point-foot, zero-moment simplification is not a limit of the method itself; robots with finite feet could use the same filter with the full six-dimensional wrench."],"forward_implications":["A quadruped can handle unexpected leg collisions without force sensors, tactile skins, or terrain maps; joint encoders and a dynamics model suffice.","The collision reflex and admittance controller shorten contact time and reduce impulse (average impulse 1.15 Ns versus 2.134 Ns in the reported comparison), and force feedback into the MPC improves balance enough to prevent a real robot from tipping.","Because the mode probabilities come from a Markov chain rather than a gait schedule, the estimator transfers across gait patterns; the paper demonstrates this on a trotting gait without retuning for swing timing.","The filter simultaneously supplies detection (mode probability) and estimation (force magnitude and direction), so downstream controllers get both signals from one estimator instead of a hierarchical thresholding pipeline."],"supporting_citations":[{"why":"Supplies the interacting multiple-model algorithm that performs mode mixing, filtering, probability update, and combination.","marker":"[21]"},{"why":"Contributes the generalized-momentum Kalman filter with disturbance observer that the mode-dependent process model is built on.","marker":"[22]"},{"why":"Provides the momentum-based observer baseline and the generalized-momentum dynamics used throughout.","marker":"[13]"},{"why":"Introduces the friction-cone and simplified pseudo-force heuristics that the mode-dependent measurement models build upon.","marker":"[10]"},{"why":"Provides the convex MPC and operational-space control framework that the force-adaptive controllers modify and are compared against.","marker":"[1]"},{"why":"A prior multiple-model Kalman filtering approach for simultaneous state estimation and contact detection that this work extends to collision mode.","marker":"[26]"},{"why":"Supplies the collision-force magnitude error metric and a robust thresholding baseline for detection comparisons.","marker":"[18]"},{"why":"Supports the assumption that one foot's collision force has negligible effect on the floating base and other legs.","marker":"[16]"}],"fun_headline_variants":["One IMM-KF detects collisions and estimates external forces","Mode-switching filter finds collisions and forces from joint data","Encoders and dynamics alone unmask leg collisions and impact forces","Collision and force estimation via a single mode-conditioned filter","Gait-independent Kalman filter spots collisions and rebalances"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an unexpected collision produces a contact force pointing roughly horizontally (inside a horizontal friction cone), while an intentional stance produces a vertical force, so the filter only needs rough prior knowledge of the terrain shape to tell them apart.","fun_headline_variants_meta":{"raw":{"variants":["One IMM-KF detects collisions and estimates external forces","Mode-switching filter finds collisions and forces from joint data","Encoders and dynamics alone unmask leg collisions and impact forces","Collision and force estimation via a single mode-conditioned filter","Gait-independent Kalman filter spots collisions and rebalances"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1507,"prompt_tokens":967,"completion_tokens":540,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":458}},"tokens_in":583,"tokens_out":540,"duration_ms":5370,"temperature":1.0,"reasoning_tokens":458,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:47:15.931499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test is to command the swing foot to strike the top of a low step while the foot is still moving downward, so the contact force is mostly vertical. If the IMM-KF assigns that event high stance probability rather than collision probability and no reflex triggers, the central detection claim fails for that case; this can be checked with a force-sensor-instrumented foot or a simulation with ground-truth contact labels.","supporting_citations":[{"cited_title":"Collision detection and identification for a legged manipulator,","cited_arxiv_id":null,"evidence_quote":"Supplies the collision-force magnitude error metric and a robust thresholding baseline for detection comparisons."},{"cited_title":"The interacting multiple model algorithm for systems with markovian switching coefficients,","cited_arxiv_id":null,"evidence_quote":"Supplies the interacting multiple-model algorithm that performs mode mixing, filtering, probability update, and combination."},{"cited_title":"Carte- sian contact force estimation for robotic manipulators using kalman filters and the generalized momentum,","cited_arxiv_id":null,"evidence_quote":"Contributes the generalized-momentum Kalman filter with disturbance observer that the mode-dependent process model is built on."},{"cited_title":"Collision detection and safe reaction with the dlr-iii lightweight manipulator arm,","cited_arxiv_id":null,"evidence_quote":"Provides the momentum-based observer baseline and the generalized-momentum dynamics used throughout."},{"cited_title":"Local reflex generation for obstacle negotiation in quadrupedal locomotion,","cited_arxiv_id":null,"evidence_quote":"Introduces the friction-cone and simplified pseudo-force heuristics that the mode-dependent measurement models build upon."},{"cited_title":"Implementation of a gait phase informed sensor- less collision detector for legged robots,","cited_arxiv_id":null,"evidence_quote":"Supports the assumption that one foot's collision force has negligible effect on the floating base and other legs."}],"review_version":1}