{"id":"fdfaafd3-b776-4e3a-abb4-12a419d544cb","arxiv_id":"2505.02414","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Some active spine control strategies make a simulated quadruped look more natural to human viewers, but all use more energy than a rigid spine.","lead":"Quadruped robots with flexible spines can look more natural to people, but they use more energy than stiff-bodied robots. Two of four spine control strategies were rated more natural than a fixed spine in a simulation study, though none saved energy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The naturalness claim rests on mean scores with no inferential statistics, and the optimised strategies' parameters were manually filtered for natural appearance, so the observed ranking may reflect sampling noise or tuning effort rather than control strategy.","rationale":"The reader's CONDITIONAL verdict already identifies the missing inferential statistics and the tuning-effort confound, so my stress-test does not move the verdict. I agree that the tuning-effort concern is real and weighty: Section 4.5 explicitly states that optimised candidates were manually chosen for natural appearance, while the real-dog strategy was not optimised at all. That asymmetry makes it impossible to attribute the naturalness ranking solely to the control strategy. However, I place the more load-bearing concern one level earlier: the paper reports no significance tests or confidence intervals for the perception scores, so even the raw claim that the two strategies 'were perceived as more natural' is not yet demonstrated. A permutation test on the existing forced-choice data would settle whether the observed mean differences are distinguishable from chance. The energy-efficiency result is a simulation finding with no error bars, but it is consistent with the plotted curves and does not undermine the paper's main claim as directly. The paper is reproducible, clearly written, and the public code and data are a genuine strength; the issues are methodological and fixable, hence CONDITIONAL remains appropriate rather than ACCEPT or REJECT.","tokens_in":15967,"tokens_out":3864,"duration_ms":54371,"concrete_test":"Using the public GitHub data, reconstruct the per-trial votes from the 49 participants and run a permutation or bootstrap test comparing each strategy's naturalness score against the fixed-spine baseline, applying a multiple-comparison correction across the four strategies. If the corrected confidence intervals for optimised-time and foot-tracking include zero, the headline claim is unsupported at the statistical level. If they exclude zero, re-run the subjective experiment with parameters selected by the same automatic cost function alone, without manual naturalness filtering, to determine whether the advantage persists when tuning effort is equalised across strategies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central positive claim is that the optimised-time and foot-tracking spine strategies were perceived as more natural than the fixed-spine baseline. This claim is not statistically established. Section 4.5 reports that the optimised parameters were found by grid search and then 'candidates with good scores were reviewed manually and chosen based on how natural they appeared to the researchers', while Section 4.9 describes a forced-choice experiment with 49 participants. Section 5.2 reports only mean naturalness scores (Figure 13) with no variance, confidence intervals, or significance tests. Without these, the observed ranking could be within sampling noise, especially when four strategies are compared against baseline. Additionally, the manual naturalness-based selection is a direct confound: the optimised-time and foot-tracking strategies received extra tuning effort specifically aimed at natural appearance, whereas the real-dog strategy used un-optimised published data and the fixed-spine baseline received no naturalness tuning. Thus the comparison does not isolate the control law from the amount of tuning effort invested in it. Both issues bear directly on the strongest claim and are addressable with the public data and a small additional experiment.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates spine control strategies for a simulated quadruped robot with a four-degree-of-freedom active spine. Four trajectory-generation strategies are compared with a fixed-spine baseline: a varying-stiffness strategy, a foot-tracking strategy, a real-dog time-varying strategy, and an optimised time-varying strategy. The authors evaluate cost of transport (CoT) for walking, trotting, and turning gaits, and they run a forced-choice online video study (49 participants) in which viewers rank naturalness and inclination to interact. The paper reports that no spine strategy improves CoT over the fixed-spine baseline, and that the optimised-time and foot-tracking strategies are perceived as more natural than the baseline. The authors conclude that spine-enabled robots may be promising for HRI applications where naturalness outweighs energy efficiency.","tokens_in":16154,"tokens_out":7812,"duration_ms":92037,"significance":"The paper addresses a genuine gap in the quadruped-spine literature, which has largely focused on efficiency and stability, by making human perception the outcome of interest. Its strengths include a transparent description of the four control strategies, a within-simulation comparison on multiple gaits, and especially the public release of code, videos, and raw data on GitHub, which allows the main analysis to be reproduced and extended. The honest reporting of a null CoT result is also valuable: it sets a boundary condition for claims that active spines improve efficiency. However, the central naturalness finding is not currently established. The reported mean scores are not accompanied by inferential statistics, and the parameter-selection procedure appears to incorporate the researchers' own judgment of naturalness, so the comparison risks confounding strategy with tuning effort. Because these two issues bear directly on the paper's main contribution, they must be addressed before the perceptual claim can be accepted.","major_comments":[{"comment":"The paper's central positive claim, that the optimised-time and foot-tracking spine strategies are perceived as more natural than the fixed-spine baseline, is supported only by point estimates. Section 5.2 reports mean naturalness scores in Figure 13, with no confidence intervals, standard errors, significance tests, or effect sizes. The experiment has a repeated-measures structure (49 participants, pairwise forced choices among five strategies), so the ranking could plausibly lie within sampling noise, especially because multiple pairwise comparisons are implicitly being made. Please provide inferential statistics for the baseline comparisons, for example a mixed-effects logistic regression on the forced-choice responses or paired nonparametric tests with multiplicity correction, and report effect sizes. The raw data are public, so these analyses should be straightforward.","section":"Section 5.2, Figure 13"},{"comment":"The naturalness ranking is confounded by unequal tuning effort across conditions. The grid search for the optimised strategies minimises the cost function in Equation (15), which does not contain a naturalness term, but the authors then state that candidates with good scores were reviewed manually and chosen based on how natural they appeared to the researchers. Thus the optimised-time strategy, and by the same procedure the foot-tracking and stiffness strategies, were selected partly for natural appearance, whereas the real-dog strategy was not optimised and the fixed-spine baseline received no naturalness-based selection. Any observed advantage for the optimised strategies could therefore reflect the manual selection and increased tuning effort rather than the control law itself. To make the claim strategy-level, the authors should either use a pre-registered, naturalness-blind selection criterion, or show that the ranking is stable across the set of near-optimal parameter candidates, or explicitly restrict their conclusion to the particular tuned instances studied.","section":"Section 4.5"},{"comment":"Hypothesis #1 is evaluated by counting how often a participant's most natural and most likely to interact with choices agree (862 of 882 votes, 97.7%), but this is not strong evidence of a correlation between the two constructs. The two questions are answered on the same videos within the same forced-choice task, so agreement is inflated by shared response tendencies and task framing, and no chance baseline or appropriate statistical model is reported. Please analyse the two responses jointly, for example with a contingency-table or multilevel model that accounts for the forced-choice design, and temper the greater likeability wording in the abstract accordingly.","section":"Section 6.1"}],"minor_comments":[{"comment":"The abstract states that the randomised trial used 50 participants, while Section 5.2 reports 50 recruited with one dropped, leaving 49 total results; please make the participant count consistent throughout.","section":"Abstract and Section 5.2"},{"comment":"The entry 4rd appears in the time-real row of the table; it should be 4th.","section":"Table 6"},{"comment":"The naturalness scoring algorithm is described only in the caption of Figure 8. Please move the full algorithm into the main text so that the normalisation and the baseline-halving rule are reproducible without reading the figure.","section":"Section 4.9, Figure 8"},{"comment":"The perception study uses videos of the simulation, and Section 4.6 explains that the MPC approximates the robot as a single rigid body with small spine movements. Please state explicitly in the limitations that the naturalness results have not been validated on hardware with true spine dynamics, since this could affect the HRI conclusions.","section":"Section 4.6 and Section 7"}],"recommendation":"major_revision","confidential_remarks":"The core idea is sound and the open-data practices are good, but the main perceptual claim needs the additional statistical treatment described above, and the tuning-effort confound must be addressed explicitly. I would not reject at this stage, but I would require either a re-analysis of the released data or a clearly narrowed claim before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere is my read on arXiv:2505.02414. It is worth knowing about because it is the first to put quadruped spine control strategies side by side in a human perception experiment, and the finding that perceived naturalness and energy efficiency can decouple is useful for HRI design. The simulation work is careful: a 4-DoF spine, MPC, four strategies including one derived from real dog data, and they release code and data.\n\nThe strong part is the study design: forced-choice comparisons against a fixed-spine baseline, randomized order, 49 participants, and a transparent scoring rule. I also credit the authors for reporting the negative energy result (no spine strategy beats fixed spine) rather than burying it.\n\nThe soft spots are real and both land on the main claim. First, there are no inferential statistics in the subjective results: Figure 13 shows mean scores with no variance, no confidence intervals, no tests. With four strategies plus a baseline, the observed ranking could easily be noise. Second, and more serious, the 'optimised time' strategy's parameters were manually filtered for natural appearance (Section 4.5), whereas the real-dog strategy was not tuned for naturalness. That means the naturalness ranking does not isolate the control law; it confounds strategy with tuning effort. The stress-test note is correct on both counts. The authors acknowledge parameter tuning difficulty in Section 4.8, but they don't address the selection bias.\n\nThat said, these are addressable, not fatal. The raw data and videos are public, so a referee could ask for a re-analysis with proper statistics and an effort-equality check. The paper is honest and clearly written; I don't see overclaiming beyond what the data as presented supports. The central claim is plausible but not yet proven.\n\nWho is this for? Researchers in quadruped control with an interest in HRI and naturalness. It deserves peer review, but with the expectation that the perception analysis be substantially strengthened before publication.","headline":"First head-to-head perception comparison of quadruped spine strategies, but the naturalness ranking rests on mean scores with no inferential statistics and a tuning-effort confound.","tokens_in":16737,"tokens_out":1845,"would_cite":false,"duration_ms":22157,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Giving a quadruped robot an actively controlled spine can make its gait look more natural to human viewers, but none of the tested spine strategies improved energy efficiency over a fixed spine.","keywords":["quadruped robot","active spine","gait naturalness","human-robot interaction","model predictive control","cost of transport","central pattern generator","bio-inspired robotics"],"falsifier":"Run a new human study that fixes tuning effort: give the real-dog law the same number of optimisation evaluations as the optimised-time law, then check whether the naturalness ranking inverts. Alternatively, measure the spine kinematics of toy-poodle-sized dogs directly; if the real-dog coefficients actually match small-dog motion and viewers still rate it least natural, the premise that real canine data looks natural on a small robot is falsified.","tokens_in":15711,"feed_emoji":"🤖","tokens_out":4083,"duration_ms":46385,"temperature":0.7,"pith_summary":"The paper asks whether adding an actively controlled spine to a quadruped robot can make its gait appear more natural to people, and whether any such gain comes at an energy cost. Using a simulated toy-poodle-sized robot with two pitch and two yaw spine joints, the authors test four spine trajectory strategies against a fixed-spine baseline across walking, trotting, and turning. In a forced-choice video study with 49 participants, the optimised time-varying and foot-tracking strategies were rated more natural than the baseline, while the real-dog time-varying and stiffness strategies scored worse. At the same time, the fixed spine had the lowest cost of transport at every tested speed, even when spine motor power was excluded. The authors conclude that perceived naturalness and energy efficiency can diverge, and that for social-robot applications like elder care, natural motion may matter more than efficiency.","feed_headline":"A moving spine makes robot gaits look natural, not efficient","feed_subtitle":"In a 49-viewer test, optimised and foot-tracking spines beat the rigid chassis on naturalness; none beat it on energy.","key_machinery":"The load-bearing components are the four spine trajectory strategies and the model predictive control (MPC) system they plug into. The spine is a four-degree-of-freedom active joint set (two pitch, two yaw) on a simulated poodle-sized robot; a central pattern generator (CPG) produces a phase variable $\\phi$ that synchronises leg swing and stance cycles with spine commands. The strategies are: a stiffness strategy that sinusoidally varies the PD gain $K_p$ while holding a fixed setpoint; a foot-tracking strategy that derives spine pitch from the average fore-aft foot displacement from neutral and spine yaw from the angle between paired feet; a real-dog time-varying strategy using a bi-periodic sine in pitch and a mono-periodic sine in yaw with coefficients taken from canine motion data; and an optimised time strategy using the same sinusoidal law with coefficients found by grid search over a weighted cost function of energy, tracking error, foot-force variance, and spine range of motion. Ground-reaction forces are computed by a representation-free MPC that treats the robot as a single rigid body, so spine motion enters as a disturbance the controller must accommodate. This pairing of an actively moved spine with an MPC that ignores spine inertia is what makes the comparison possible, and it shapes the energy results.","core_discovery":"The paper claims that spine motion in a quadruped robot is judged more natural when it is subtle and coordinated with the legs, and that copying real canine motion does not automatically look natural on a small robot. Concretely, in a randomised comparison with 49 participants, the optimised time-varying strategy and the foot-tracking strategy scored higher than the fixed-spine baseline, while the real-dog time-varying strategy and the stiffness strategy scored lower. The same participants chose the same gait as most natural and most appealing to interact with in 97.7% of votes. Dynamic measurements show the fixed spine has the lowest cost of transport at all tested velocities, and spine strategies increase the work done by the legs; the two strategies judged most natural also showed the most consistent trotting footfalls, suggesting a possible link between perceived naturalness and gait regularity rather than energy cost.","pith_inferences":["The paper's explanation for the real-dog strategy's low naturalness (viewers compared the small simulated robot to larger dogs) is testable: re-render the same gaits with a familiar size reference object in view, and the naturalness ranking may shift.","The observed correlation between naturalness and footfall consistency in trotting could be isolated experimentally: if the spine is hidden and only footfall timing is varied, one could test whether regularity alone drives perceived naturalness.","The MPC treats spine inertia as negligible, so the energy verdict may be specific to this controller; modelling spine dynamics inside the MPC, as the paper suggests, could change both cost of transport and perceived naturalness.","The naturalness-versus-efficiency trade-off may not hold on larger robots; on a platform closer in scale to the dogs in the motion-capture data, the real-dog coefficients might score higher and the conclusions could invert."],"forward_implications":["For human-robot interaction applications such as elder-care companions, a spine that moves subtly and tracks the feet can make a quadruped seem more natural and more worth interacting with, even if it does not save energy.","Gait naturalness and energy efficiency are not coupled: the most efficient gait (fixed spine) was not the most natural, and no spine strategy improved cost of transport over the baseline.","Robot designers should select the spine strategy per gait: the stiffness strategy jumped to most natural in trotting, while the optimised-time strategy was best in walking and turning, implying that HRI-focused robots should be able to switch strategy.","The 97.7% match between 'most natural' and 'most likely to interact with' suggests that perceived naturalness is a reliable proxy for user acceptance in this context.","Because the fixed spine was most efficient even with spine motor power excluded, active spine motion increases leg work, so an energy-neutral natural spine would likely need passive or hybrid components."],"supporting_citations":[{"why":"Supplies the canine spine motion data (Wachs et al.) used to set the real-dog time-varying strategy's sine coefficients.","marker":"[9]"},{"why":"Provides the representation-free MPC controller that all spine strategies are tested with, treating the robot as a single rigid body.","marker":"[27]"},{"why":"Shows the sinusoidal spine-leg coordination approach that the time-varying strategies extend, and reports prior energy-efficiency gains from spine actuation that this paper's results are compared against.","marker":"[19]"},{"why":"One of the variable-stiffness spine works (with [17] and [18]) that the stiffness strategy is based on.","marker":"[16]"},{"why":"Shows a musculoskeletal robot with variable trunk stiffness, supporting the stiffness-variation approach used in one of the four strategies.","marker":"[17]"},{"why":"Describes a quadruped bionic robot with a variable-stiffness spine, another reference for the stiffness strategy.","marker":"[18]"},{"why":"An early biomimetic quadruped with an active spine, referenced as a foot-tracking precedent along with [23], [26], and [44].","marker":"[20]"},{"why":"Analyses an active artificial spine in a quadruped robot, reporting leg power reduction that this paper's foot-tracking results directly contradict.","marker":"[23]"},{"why":"Hildebrand's symmetric gait classification, used to produce the footfall-consistency plots for the ethological analysis.","marker":"[54]"}],"fun_headline_variants":["Copying real dog spine motion makes robots look less natural","Robot spine: subtle coordinated motion wins over dog mimicry","Natural gait beats energy efficiency in quadruped robot spine study","Human viewers judge subtle spine motion more natural than dog-like","Spine strategy tradeoff: naturalness up, energy efficiency down"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The strategies are compared under equally well-tuned parameters: the optimised-time strategy received a grid search plus manual review by the researchers, while the real-dog strategy used manually chosen, unoptimised coefficients, so the naturalness ranking could reflect tuning effort rather than the control law itself.","fun_headline_variants_meta":{"raw":{"variants":["Copying real dog spine motion makes robots look less natural","Robot spine: subtle coordinated motion wins over dog mimicry","Natural gait beats energy efficiency in quadruped robot spine study","Human viewers judge subtle spine motion more natural than dog-like","Spine strategy tradeoff: naturalness up, energy efficiency down"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000362,"raw_usage":{"total_tokens":1949,"prompt_tokens":934,"completion_tokens":1015,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":932}},"tokens_in":550,"tokens_out":1015,"duration_ms":9329,"temperature":1.0,"reasoning_tokens":932,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:51:09.049922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a new human study that fixes tuning effort: give the real-dog law the same number of optimisation evaluations as the optimised-time law, then check whether the naturalness ranking inverts. Alternatively, measure the spine kinematics of toy-poodle-sized dogs directly; if the real-dog coefficients actually match small-dog motion and viewers still rate it least natural, the premise that real canine data looks natural on a small robot is falsified.","supporting_citations":[{"cited_title":"Three-dimensional movements of the pelvis and the lumbar intervertebral joints in walking and trotting dogs","cited_arxiv_id":null,"evidence_quote":"Supplies the canine spine motion data (Wachs et al.) used to set the real-dog time-varying strategy's sine coefficients."},{"cited_title":"Representation-free model predictive control for dynamic motions in quadrupeds","cited_arxiv_id":null,"evidence_quote":"Provides the representation-free MPC controller that all spine strategies are tested with, treating the robot as a single rigid body."},{"cited_title":"High speed trot-running: Implementation of a hierarchical controller using proprioceptive impedance control on the mit cheetah","cited_arxiv_id":null,"evidence_quote":"Shows the sinusoidal spine-leg coordination approach that the time-varying strategies extend, and reports prior energy-efficiency gains from spine actuation that this paper's results are compared against."},{"cited_title":"Realization of stable quadruped gait transition by chang- ing body stiffness (japanese)","cited_arxiv_id":null,"evidence_quote":"One of the variable-stiffness spine works (with [17] and [18]) that the stiffness strategy is based on."},{"cited_title":"A study on trunk stiffness and gait stability in quadrupedal locomotion using musculoskeletal robot","cited_arxiv_id":null,"evidence_quote":"Shows a musculoskeletal robot with variable trunk stiffness, supporting the stiffness-variation approach used in one of the four strategies."},{"cited_title":"Design of quadruped bionic robot with variable stiffness spine","cited_arxiv_id":null,"evidence_quote":"Describes a quadruped bionic robot with a variable-stiffness spine, another reference for the stiffness strategy."},{"cited_title":"Design and development of biomimetic quadruped robot for behavior studies of rats and mice","cited_arxiv_id":null,"evidence_quote":"An early biomimetic quadruped with an active spine, referenced as a foot-tracking precedent along with [23], [26], and [44]."},{"cited_title":"Analysis of using an active artificial spine in a quadruped robot","cited_arxiv_id":null,"evidence_quote":"Analyses an active artificial spine in a quadruped robot, reporting leg power reduction that this paper's foot-tracking results directly contradict."}],"review_version":1}