{"id":"b4c533e6-07ff-4ab9-b3f5-9271283e6140","arxiv_id":"2509.09842","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On-field test of a five-sensor headband against an instrumented mouthpiece finds good time-history agreement for angular velocity and translational acceleration, weaker agreement for angular acceleration, and a 40.9% normalized peak angular velocity bias.","lead":"Researchers tested a sensor-equipped headband against an instrumented mouthpiece during 18 soccer headers on one player. The headband agreed well on angular velocity and translational acceleration time histories, but showed large peak bias, especially for angular velocity.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Mouthpiece reference validity is the load-bearing assumption: if the Wake Forest mouthpiece is biased, headband–mouthpiece agreement cannot establish headband accuracy.","rationale":"I read the full manuscript in good faith. The paper is careful and modest: it uses 18 headers from one subject, reports CORA scores ranging from 'fair' to 'excellent', and explicitly discloses the large peak angular velocity bias, the small sample, the unquantified headband fit/slip, and the mouthpiece limitations. The central claim—'reasonable agreement with the mouthpiece for some kinematic measures and impact conditions'—is scoped and does not overclaim interchangeability. However, the entire evaluation is framed as a field validation of the headband against a widely accepted reference. If that reference is biased, the agreement numbers are uninterpretable as validation of the headband's accuracy. The authors themselves acknowledge this in §4.5, noting the mouthpiece has not been cadaver-validated and that mandible motion can cause errors; they suggest future studies include 3D high-speed video. This is the most load-bearing concern because it is presupposed by every reported comparison and is not an internal inconsistency but an external-validity risk. The concrete test I propose—independent optical validation of the mouthpiece across the exact kinematics range studied—would settle whether the reference error is small enough to trust the headband–mouthpiece agreement. Given the paper's own caveats and conditional language, the existing CONDITIONAL verdict remains appropriate; no change is needed. I agree with the reader's weakest_assumption.","tokens_in":16476,"tokens_out":5497,"duration_ms":69811,"concrete_test":"Conduct a controlled laboratory experiment on a cadaver head (or an ATD with realistic jaw/mandible coupling) with the Wake Forest mouthpiece mounted together with a research-grade optical motion capture system (≥1 kHz, multi-camera). Deliver impacts spanning the reported field ranges (ball speeds 32–38 km/h; mouthpiece PRV 4–10 rad/s, PRA 610–3200 rad/s², PLA 120–340 m/s²). Compute mouthpiece-vs-optical peak bias and CORA scores for angular velocity, angular acceleration, and translational acceleration. If the mouthpiece bias exceeds ~10% or CORA falls below 'good' (>0.66) on any of these metrics, then the headband–mouthpiece agreement cannot by itself support field accuracy and the conclusion should be weakened to 'device-to-device agreement only.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the headband showed reasonable agreement with the mouthpiece, and the surrounding framing treats this as field validation of the headband for measuring skull kinematics. The load-bearing assumption is that the Wake Forest instrumented mouthpiece is an accurate reference for skull kinematics. The authors explicitly concede in §4.5 that mouthpieces are not ground truth, that mandible motion can bias them, and that this particular retainer-form mouthpiece has been validated only on a clenched-mandible ATD, not on a cadaver. If the mouthpiece's own error is comparable to or larger than the observed headband–mouthpiece differences (mean bias 40.9% for angular velocity, 16.6% for translational acceleration, −14.1% for angular acceleration), then the reported CORA and Bland-Altman agreement cannot be interpreted as evidence that the headband measures skull kinematics in the field. This concern is structurally distinct from sampling variability: it would invalidate the comparison even with a larger sample because the reference itself is unverified. The manuscript provides no quantitative bound on the reference error, so the central validation claim rests on an uncharacterized reference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a field evaluation of a five-IMU instrumented headband against a custom instrumented mouthpiece (Wake Forest) on one adult female soccer player performing 18 controlled headers (six throw-ins, six goal-kicks, six corner-kicks). Headband angular velocity is reconstructed by averaging five gyroscopes and adaptive wavelet filtering; angular and translational accelerations are reconstructed using finite differentiation and an algebraic A3G1 algorithm. Agreement is quantified by CORA time-history scores and Bland-Altman peak-bias analyses. The authors report 'good' average CORA scores (angular velocity 0.79±0.08, translational acceleration 0.73±0.05, angular acceleration 0.67±0.06 with A3G1 and 0.62±0.08 with differentiation), but large peak biases: 40.9% normalized mean bias for angular velocity, 16.6% for translational acceleration, and -14.1% (A3G1) or -4.27% (differentiation) for angular acceleration. They conclude that the headband shows reasonable agreement with the mouthpiece for some kinematic measures and impact conditions.","tokens_in":16785,"tokens_out":6361,"duration_ms":67975,"significance":"If the comparison is accepted, this is a useful contribution to the sparse in-vivo validation literature for headband-type sensors. The study is methodical: the processing chain is described in detail, CORA uses recommended parameters, Bland-Altman analyses are standard, and the authors explicitly list limitations (single subject, small sample, mouthpiece not ground truth). The direct comparison with the Wu et al. skull-cap/skin-patch data is informative, and the paper follows CHAMP reporting guidelines. The main result—good time-history agreement but non-negligible peak bias—is plausible and actionable for sensor development. However, the strength of the conclusion depends on an uncharacterized reference instrument and on interpreting 'reasonable agreement' in light of the large bias percentages; these issues need to be addressed before the paper can be accepted as a field validation rather than a single-subject feasibility study.","major_comments":[{"comment":"The paper concedes that mouthpieces are not ground truth, that mandible motion can bias them, and that this specific retainer mouthpiece has been validated only on a clenched-mandible ATD, not a cadaver. This is directly load-bearing because the abstract's conclusion ('reasonable agreement with the mouthpiece') and the discussion's language ('evaluate the headband in measuring the full head kinematics') treat mouthpiece readings as the reference. Without a quantitative bound on reference error, the 40.9% PRV bias could be entirely a mouthpiece artifact, or the true headband error could be larger than reported. I recommend either (a) explicitly reframing the paper as an agreement study between two wearable devices, or (b) adding a quantitative uncertainty analysis using published mouthguard validation data to bound the reference error. The current version overreaches for a validation clai","section":"§4.5"},{"comment":"The 18 headers are clustered within one subject and one fitting. The paired t-tests applied to header-type subgroups (n=6 each) treat repeated trials as independent, which inflates significance and narrows error bars. The authors correctly acknowledge the sample-size limitation in §4.5, but the inferential statements in §4.2 (e.g., 'not statistically significant') and the header-type comparisons in Fig. 8 should be either removed or replaced with a repeated-measures analysis that accounts for within-subject correlation. At minimum, the paper should state that all p-values are descriptive and not corrected for clustering.","section":"§2.5, §3, Fig. 8"},{"comment":"The 'reasonable agreement' verdict rests mainly on CORA scores, but CORA is a weighted combination of phase, magnitude, and shape. The Bland-Altman analysis shows a normalized mean bias of 40.9% (3.85 rad/s) for peak angular velocity, with limits of agreement spanning roughly -0.3 to 8.0 rad/s. For the metric most often associated with injury risk, this is a substantial error. The manuscript should state an explicit acceptance threshold or cite one from the head-impact-sensor literature, and discuss how a 40.9% PRV bias would affect brain-strain or injury-risk estimates. Without this, the qualitative conclusion is underdetermined.","section":"§3, Fig. 7a, Abstract"},{"comment":"The mouthpiece reference angular acceleration is obtained by numerical differentiation of its angular velocity (five-point stencil), while the headband A3G1 result is computed algebraically from accelerometer/gyroscope data. The comparison is therefore asymmetric: the reference itself contains differentiation-amplified noise. The paper notes this in passing, but the CORA and bias results for angular acceleration should also be reported using a common processing path (e.g., differentiating both signals) to separate algorithmic differences from reference-processing artifacts. This is especially important because the A3G1 method is a central novel component.","section":"§2.2, §2.4.2, §4.2"}],"minor_comments":[{"comment":"Equation (3) is referred to as 'Eq. 2.4.2' in the text; fix the cross-reference.","section":"§2.4.3"},{"comment":"The normalized Bland-Altman bias is computed relative to the maximum mouthpiece reading. Please also report the bias relative to the mean of the paired measurements, and justify the denominator choice; normalization by max can inflate or deflate percentages depending on impact severity.","section":"§2.5, Fig. 8b"},{"comment":"The sensitivity analysis for the t=150 ms end point is only in Supplementary Fig. S1. Please summarize the result in the main text, since the cutoff frequency f0 depends on this choice.","section":"§2.4.1"},{"comment":"The NRMS comparison with Wu et al. uses different window lengths (24.4 ms vs. the present study's window). State the window length used for the headband NRMS values and confirm the comparison is apples-to-apples.","section":"§4.4"},{"comment":"p-values are mentioned but not reported for the header-type comparisons; give exact values or confidence intervals in the text or figure.","section":"§4.2, Fig. 8"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and within scope. The main risk is overclaiming field validation from a single-subject comparison against an unvalidated-in-cadaver mouthpiece. I would recommend the editors ask for a revision that either narrows the claim to 'agreement with mouthpiece' or provides a quantitative reference-uncertainty analysis. No concerns about author conduct."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the field data: 18 real soccer headers on a human subject, with simultaneous headband and instrumented mouthpiece recordings. That is a legitimate extension of the authors' lab work, and it answers a real gap, because most on-field wearable evaluations only check event detection, not kinematic waveforms.\n\nWhat the paper does well: the methods are described in enough detail to be followed, the metrics (CORA, Bland-Altman, NRMS) are standard, and the claims are carefully worded. The abstract says “reasonable agreement for some kinematic measures and impact conditions,” which is exactly what the data show. The authors also do an honest job in Section 4.5 listing limitations: single subject, small n, mouthpiece not ground truth, possible mandible effects, no cadaver validation, and the need for high-speed video in future work. They do not overfit the reference; the wavelet filter and A3G1 algorithm come from prior work and are not tuned to the mouthpiece data.\n\nThe soft spots are real but not fatal. One subject and 18 headers mean the point estimates are shaky, and the 40.9% normalized bias on peak angular velocity is large enough to stop anyone from using the headband interchangeably with a mouthpiece. The stress-test concern about the mouthpiece reference is valid: if the Wake Forest device is biased, the headband's apparent errors are miscalibrated. But the paper already concedes this, and the central claim is about agreement between two devices, not about absolute skull kinematics. What is missing is any quantitative bound on the reference uncertainty, and that should be added or at least discussed more explicitly. The A3G1 angular accelerations miss the earliest peaks, so the better CORA score for that method is partly misleading.\n\nI also note no data or code are shared, which limits reproducibility, though this is common for preliminary field studies.\n\nOverall, this is a solid preliminary study with modest claims. It deserves peer review, not desk rejection. The right outcome is likely revision with more subjects, more impacts, and a sharper treatment of reference-device uncertainty. I would not cite it as a definitive validation, but I would point to it as the first field comparison for this headband.","headline":"A modest, clearly-reported field validation of a previously developed headband, with the expected caveat that the mouthpiece reference is not ground truth and the sample is one subject.","tokens_in":17246,"tokens_out":1243,"would_cite":true,"duration_ms":16285,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a five-sensor instrumented headband, previously validated in the lab, can capture some head kinematics of soccer headers under real field conditions, with time-history agreement ranging from 'good' to 'excellent' for","keywords":["mild traumatic brain injury","instrumented headband","soccer headers","field validation","head kinematics","angular velocity","translational acceleration","wearable sensors"],"falsifier":"Run the same 18-header protocol with the athlete also tracked by high-speed biplanar radiography or a skull-pin-mounted reference; if the mouthpiece deviates from that reference by as much as the headband deviates from the mouthpiece, the reported headband agreement is not established.","tokens_in":16432,"feed_emoji":"⚽","tokens_out":4904,"duration_ms":52469,"temperature":0.7,"pith_summary":"The paper evaluates an instrumented headband worn by a human soccer player heading balls from throw-ins, goal-kicks, and corner-kicks, comparing its reconstructed head kinematics to those from an instrumented mouthpiece worn simultaneously. Time-history agreement (CORA scores) was 'good' to 'excellent' for angular velocity (0.79±0.08) and 'good' for translational acceleration (0.73±0.05), but only 'fair' to 'good' for angular acceleration (0.62–0.67). Peak-value agreement was weaker: the headband over-predicted peak angular velocity by 40.9% and peak linear acceleration by 16.6% on average, and under-predicted peak angular acceleration by 14.1%. The paper concludes that the headband can be used in field studies for selected kinematic measures and impact conditions, but not interchangeably with a mouthpiece.","feed_headline":"Headband matches mouthpiece on 2 of 3 head-motion measures","feed_subtitle":"Angular velocity and linear acceleration waveforms agree well; peak velocity bias still hits 41 percent.","key_machinery":"The headband integrates five triaxial inertial measurement units around the occipital region. Angular velocity is reconstructed by averaging the five gyroscope signals and applying a continuous-wavelet-transform-based adaptive filter that selects a per-impact cutoff frequency to remove transient sensor noise while preserving the true signal. Angular and translational acceleration are computed either by differentiating the filtered angular velocity or by the A3G1 algorithm, which solves a rigid-body algebraic system using three accelerometers and one gyroscope to avoid noise amplification from differentiation. Agreement with the mouthpiece is quantified with CORA scores for time histories and","core_discovery":"The central claim is that the instrumented headband, when evaluated on a human soccer player under realistic field conditions, achieves good time-history agreement with a reference instrumented mouthpiece for angular velocity (CORA = 0.79±0.08) and translational acceleration (CORA = 0.73±0.05), while angular acceleration agreement is lower (0.67±0.06 with an algebraic A3G1 method, 0.62±0.08 with numerical differentiation). Peak kinematics, however, show substantial bias: mean bias reached 40.9% of the maximum mouthpiece reading for angular velocity, 16.6% for translational acceleration, and -14.1% for angular acceleration. The paper argues the headband is suitable for field deployment when t","pith_inferences":["The 41% peak-velocity bias suggests the headband is currently unsuitable for computing brain-strain surrogates in individual impacts; even if time histories look similar, peak errors of this size will propagate nonlinearly into strain estimates.","A spring-dashpot correction model, analogous to what prior work applied to skin patches and skull caps, could be fit to the headband-mouthpiece bias and might reduce the peak errors without hardware changes.","The A3G1 algorithm's failure to capture peaks in the first ~15 ms after impact hints that the accelerometer signals used in the algebraic solve are themselves contaminated by the same transient noise the wavelet filter removes; testing the algorithm on the unfiltered gyroscope signal could isolate the error source.","Evaluating the headband against a second reference (e.g., biplanar video) on the same subject would test whether the mouthpiece or the headband is the larger source of disagreement, a question the current single-reference design cannot answer."],"forward_implications":["For large-cohort soccer studies that need angular velocity and linear acceleration time histories, the headband may be sufficient, since CORA scores above 0.7 indicate good waveform agreement.","Peak-based injury-risk metrics should not be derived from the current headband: the 40.9% peak angular-velocity bias exceeds what interchangeable-sensor studies typically accept.","The A3G1 algebraic method improves angular-acceleration time-history agreement (0.67 vs 0.62) but worsens peak bias (-445 vs -135 rad/s²), so method choice depends on whether peaks or waveforms matter more.","The lab-to-field drop in filter cutoff frequency (126±50 Hz to 56±33 Hz) shows that lab validation alone is insufficient; field-specific filtering or hardware changes are needed.","Headband fit, hair and soft-tissue coupling, and pre-impact head motion are identified as main causes of the field performance gap."],"fun_headline_variants":["Headband tracks soccer headers, but velocity bias lingers","Field test: Headband vs mouthpiece on head impacts","Wearable headband: good waveform, bad peak bias","Head kinematics: headband matches on two measures","Soccer headband shows promise, but angular accel lags"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The evaluation assumes the instrumented mouthpiece yields accurate skull kinematics; the paper itself notes that mouthpieces are not ground truth, can move with the mandible, and the specific mouthpiece has not been cadaver-validated, so a biased reference would misstate the headband's true accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Headband tracks soccer headers, but velocity bias lingers","Field test: Headband vs mouthpiece on head impacts","Wearable headband: good waveform, bad peak bias","Head kinematics: headband matches on two measures","Soccer headband shows promise, but angular accel lags"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1178,"prompt_tokens":864,"completion_tokens":314,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":232}},"tokens_in":608,"tokens_out":314,"duration_ms":4053,"temperature":1.0,"reasoning_tokens":232,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:35:53.106441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 18-header protocol with the athlete also tracked by high-speed biplanar radiography or a skull-pin-mounted reference; if the mouthpiece deviates from that reference by as much as the headband deviates from the mouthpiece, the reported headband agreement is not established.","supporting_citations":[],"review_version":1}