{"id":"b2835028-6bb6-42bf-affa-dd36c5fc2982","arxiv_id":"2505.04548","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A movable, low-cost robotic dummy head with binaural microphones and a mouth simulator enables repeatable dynamic audio recordings in lab settings.","lead":"This paper presents a low-cost 3D printed dummy head with ear microphones, a mouth speaker, and a quiet motorized turntable, so it can talk, listen, and rotate during audio experiments. The device can automate binaural recordings and, for the first time, produce repeatable recordings with a moving sound source, which could speed up audio research and dataset creation.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Motor-noise validation is indirect: Figs. 5–6 measure at 1 m, and the repeatability metrics in Figs. 9–10 cannot detect deterministic motor noise, so the quiet-motor claim is unestablished at the ear microphones.","rationale":"The reader's weakest assumption is close but not identical. The reader worried that the ear microphones are near the turntable motor; in the demonstrated moving-talker configuration the binaural microphones are on a separate stationary listener, so the actual gap is that no measurement is made at those microphones and the reported metrics are structurally blind to deterministic motor noise. This makes the concern more load-bearing: even the repeatable result would hold for a noisy but repeatable motor. The fix is a simple in-situ noise-floor measurement; because the design files and code are open, this is readily testable. The paper has real independent support: HRTF and ILD comparison to KEMAR, ITD agreement, and reproducible open-source hardware; those parts are not in question. The concern is confined to the motor-noise validation, which underpins the abstract's differentiator and the suitability claim. Verdict remains conditional: accept with the added condition that in-situ motor-noise measurements at the ear microphones be reported, or the claims be narrowed to repeatable rather than quiet.","tokens_in":8070,"tokens_out":7219,"duration_ms":73204,"concrete_test":"Run the Fig. 8 experiment with the listener head's ear microphones, and record three conditions with the talker head: (i) motor stopped, loudspeaker silent; (ii) motor rotating at 0.2 rev/s, loudspeaker silent; (iii) motor rotating, loudspeaker reproducing the target. Compute the ear-microphone power spectra of (i) and (ii) and compare them with the target level in (iii) at the same geometry. If the motor-on component exceeds the ambient noise floor by more than a few dB in any speech-relevant band, or if it is not at least 20 dB below the target level, repeatability alone is insufficient to support the quiet-motor claim. Report the talker-listener distance and orientation used.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing premise is that the moving device is acoustically unobtrusive enough for objective evaluation of audio algorithms (abstract; Section 2.2). The evidence for this is indirect in two ways. First, the only direct motor-noise measurements (Figs. 5–6) place a microphone 1.0 m from the device and apply spectral subtraction; Fig. 7 shows vibration damping, but none of these measurements are made at the ear-microphone positions used in the validation experiments. Second, the benchmark metrics would not reveal motor noise even if it were present. In Fig. 9, the physically mixed recording and the artificially mixed recording are compared; if the target recording contains motor noise from the rotating talker head, that same motor noise appears in both mixtures, so high agreement is expected regardless of its level. In Fig. 10, the target speech recording is repeated eight times; a stepper motor driven by the same control signal produces nearly deterministic, repeatable noise, so a sample-wise average that includes the motor noise will still be matched closely by each repetition. Thus repeatability and artificial/physical agreement establish that the motor noise is repeatable, not that it is quiet. If motor noise is audible at the listener's ears, the SNR-gain evaluations in Section 4.3 are biased because motor noise is folded into the target or mixture signals rather than treated as an independent interferer. The manuscript does not report an in-situ motor-on/motor-off comparison at the ear microphones, so the central claim that the device enables high-quality dynamic recordings is conditional on an untested assumption.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a low-cost, open-source robotic dummy head: a 3D-printed acoustic mannequin with binaural ear microphones and a mouth loudspeaker, mounted on a stepper-motor turntable. The authors argue that the device combines the acoustic realism of conventional mannequins with precise, quiet mobility, enabling automated spatially-stationary experiments and repeatable spatially-dynamic recordings. Validation includes HRTF comparisons against a KEMAR mannequin (qualitative magnitude plots, quantitative ILD and ITD), motor-noise measurements at 1 m distance, repeatability tests over eight repetitions, a comparison of artificially mixed versus physically mixed recordings, and a demonstration of a motion-robust binaural MVDR beamformer. The design files are provided as open source.","tokens_in":8341,"tokens_out":2694,"duration_ms":28713,"significance":"If the central claim is sustained, the device would be a useful research tool for repeatable dynamic binaural audio experiments and for generating large labeled spatial-audio datasets. The paper's strengths include its explicit external validation against KEMAR (with a reported ITD RMSE of 67.9 microseconds), the use of eight repeated recordings to demonstrate repeatability, the comparison of artificial versus physical mixing, and the release of open-source design files. These features make the claims independently checkable and are appropriate for a tools-focused audio research paper.","major_comments":[{"comment":"The 'quiet motor' claim is load-bearing for the entire paper, but it is validated only with a microphone placed 1.0 m from the device, not at the binaural ear-microphone positions used in the actual experiments. The ear microphones are mounted in the printed head close to the turntable motor, so structure-borne and near-field airborne motor noise could be substantially higher at those positions than at 1 m. The paper should report an in-situ motor-on versus motor-off measurement at the ear microphones, ideally in terms of the resulting SNR or noise spectrum relative to the target speech level, because this is the quantity that determines whether the device is acoustically unobtrusive for objective evaluation.","section":"Section 2.2, Figs. 5-7"},{"comment":"The repeatability and artificial/physical mixing experiments cannot detect deterministic motor noise, so they do not by themselves establish that the motor is quiet. In Fig. 9, if the target recording contains motor noise, that same noise appears in both the artificially mixed and physically mixed signals, so high agreement is expected regardless of the noise amplitude. In Fig. 10, repeated target recordings from the same stepper motor driven by the same control signal will contain nearly identical motor noise, and comparing each repetition to the sample-wise average will show high repeatability even if the motor noise is substantial. These results demonstrate that the motor noise is repeatable, not that it is absent or inaudible. A separate in-situ noise measurement is needed to support the quiet-motor claim.","section":"Section 3, Figs. 9-10"},{"comment":"The beamforming comparison between static and moving talkers could be confounded by motor noise, because any motor noise radiated during motion is folded into the target or mixture signals rather than modeled as an independent interferer. If motor noise is present at the ear microphones during the moving condition, the observed high-frequency SNR-gain drop could be partly attributable to that noise rather than to the effect of motion on RTF estimation. The authors should either provide the in-situ motor-noise measurement recommended above, or analyze the beamforming result with motor-off reference recordings, before interpreting the 'surprising result' as a motion-related phenomenon.","section":"Section 4.3, Fig. 11"}],"minor_comments":[{"comment":"The caption for Fig. 10 appears to be copied from Fig. 8 and describes a beamforming setup, whereas the text and the figure itself describe repeatability of repeated target speech recordings. The caption should be corrected to describe what is actually plotted.","section":"Fig. 10 caption"},{"comment":"The text reports a 'root mean-squared error (MSE) of 67.9 µs'; since the quantity is in microseconds, this should be called RMSE rather than MSE, or the abbreviation should be defined consistently.","section":"Section 2.1, Fig. 3 text"},{"comment":"The HRTF comparison is mostly qualitative; adding a quantitative frequency-dependent magnitude error metric (for example, average absolute dB difference per frequency band between the printed head and KEMAR) would strengthen the acoustic-realism claim.","section":"Fig. 2 and Fig. 3"},{"comment":"The forgetting factor α is said to correspond to a time constant τ, but the explicit relationship (such as α = exp(-1/(τ·fs)) or the equivalent per-frame formula) is not given. Please state this relationship so the parameter setting is reproducible.","section":"Section 4.2, Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful open-source research tool—a 3D-printed dummy head with a mouth simulator on a quiet direct-drive stepper turntable—and the validation mostly delivers. The one gap that matters is the motor noise claim: the only direct measurements are from a microphone 1 m away, not at the binaural ear mics. The repeatability metrics can't catch a deterministic motor tone. So the 'quiet enough at the ear' premise is plausible but unproven.\n\nWhat's good: open design files (GitHub), ITD RMSE 67.9 µs vs KEMAR, ILD and radiation comparisons, 8-repeat sample-wise error benchmark, and a sensible beamforming demo. The authors are honest that the beamforming result is preliminary. This is exactly the kind of hardware paper that should be published with code and CAD files. The incremental step over refs [9] and [25] is real: a quiet turntable that allows recording during motion, not just between motions.\n\nSoft spots, in order of importance. First, the motor-noise evidence is indirect. Fig. 5–6 place a microphone 1.0 m away and use spectral subtraction; Fig. 7 shows structure damping. None of this tells you what the ear mics hear. The stress-test note is correct that Fig. 9 (artificial vs physical mix) and Fig. 10 (8 repeats) can't separate a deterministic motor tone from the target. An in-situ motor-on/motor-off recording at the ear mics would settle it. This is a small addition, not a redesign. Second, the HRTF agreement is qualitative, but that's acceptable for a tool paper; the ITD/ILD numbers carry the quantitative load. Third, production issues: the caption to Fig. 10 is mislabeled (repeats Fig. 8) and there's a broken pdf include in the beamforming section. Those are trivial but should be fixed.\n\nWho it's for: audio signal processing groups that need repeatable dynamic binaural recordings without a full KEMAR and human actors. It won't reshape the field, but it removes a practical bottleneck. It deserves peer review—the revision should gate acceptance on the in-situ motor noise measurement. My verdict matches the reader's: conditional.","headline":"Useful open-source robotic dummy head with mostly sound validation, but the quiet-motor claim needs an in-situ ear-microphone measurement before the headline contribution is fully established.","tokens_in":8911,"tokens_out":2830,"would_cite":true,"duration_ms":26709,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 3D-printed robotic dummy head records repeatable binaural audio while moving, bringing realistic motion into objective audio evaluation.","keywords":["robotic dummy head","binaural audio","head-related transfer function","spatially-dynamic recordings","repeatable audio experiments","motor noise","MVDR beamforming","3D-printed acoustics"],"falsifier":"Place a measurement microphone at or inside the ear canal of the rotating head in a quiet room and compare spectra with the motor still and rotating at 0.2 to 0.4 rev/s; if motor harmonics exceed the stationary noise floor at speech frequencies, the claim that recordings are uncontaminated and repeatable fails for those speeds.","tokens_in":7872,"feed_emoji":"🤖","tokens_out":7440,"duration_ms":65273,"temperature":0.7,"pith_summary":"The paper argues that a low-cost, 3D-printed dummy head on a quiet turntable can move, talk, and listen during audio recordings without contaminating them with motor noise. This would close the gap between stationary acoustic mannequins, which are realistic but immobile, and mobile robots, which have been too noisy for clean recordings. The authors validate acoustic realism by comparing the head's HRTF, interaural level differences, and mouth radiation pattern with a KEMAR mannequin, and they show that artificially mixed recordings closely match physically mixed ones even while the talker rotates. Eight repeated recordings agree closely, supporting the claim that dynamic recordings are repeatable and suitable for objective evaluation of audio algorithms.","feed_headline":"Quiet robot dummy head records motion in repeatable audio tests","feed_subtitle":"A low-cost printed head on a whisper-quiet turntable makes moving-source audio experiments repeatable.","key_machinery":"The load-bearing mechanism is the quiet turntable: a gear-free, direct-drive stepper motor driven by a specialized control algorithm and operated below 0.4 revolutions per second, where motor harmonics are relatively spectrally white, with a 3D-printed structure that dampens vibration and no cooling fan. This lets the head rotate during recording without audible contamination. The other half is the 3D-printed dummy head, whose measured HRTF, interaural level and time differences, and mouth radiation pattern are qualitatively similar to a KEMAR, giving lifelike spatial cues at low cost and low mass. Together the two parts convert a binaural mannequin from a stationary measurement instrument into a repeatable moving source and listener.","core_discovery":"The central claim is that spatially-dynamic audio recordings can be made repeatable and physically realistic by using a quiet, robotically rotated acoustic dummy head. The device combines a 3D-printed head with two in-ear microphones and a mouth loudspeaker, mounted on a direct-drive stepper motor turntable whose motor noise stays at or below whisper level near conversational distances. Benchmark experiments show that the normalized mean-squared error between artificially mixed and physically mixed waveforms remains low at all tested rotation speeds, and repeated recordings are highly consistent. The authors conclude that these robot-enabled recordings are repeatable and suited for objective evaluation, and they demonstrate the utility by applying a motion-robust minimum-variance distortionless-response (MVDR) beamformer to a moving talker, finding that the presence of motion, not its rate, causes a severe drop in high-frequency SNR gain.","pith_inferences":["If self-noise at the ear microphones proves as low as the far-field measurements suggest, the same platform could generate large labeled datasets for sound localization, head-pose estimation, and cocktail-party separation with physically moving sources rather than simulated room impulse responses.","An immediate test the paper does not report is placing a probe microphone at or inside the ear position during rotation; this would directly confirm the binaural recordings are free of motor contamination at the speeds recommended for experiments.","The preliminary beamformer result suggests motion itself may disrupt relative transfer function tracking; varying the rotation trajectory or the adaptation time constant would reveal whether the high-frequency drop is fundamental or an artifact of the 200 ms forgetting factor.","The one-axis turntable could be extended to a second rotation axis or a mobile base, letting researchers approximate natural head and body motion while keeping the same repeatable recording protocol."],"forward_implications":["Researchers can automate the labor-intensive data collection for binaural and spatial audio experiments by scripting head rotations instead of repositioning loudspeakers or people.","Separate noise-only and target recordings from a moving source can be scaled and summed to synthesize arbitrary SNR conditions without extra recording passes, with error comparable to ambient noise.","Dynamic scenarios involving a rotating talker become objectively evaluable: repeated recordings with identical motion show high repeatability, which human actors cannot provide.","Because the design uses standard 3D printing and a low-cost stepper motor, laboratories can build the device and reproduce dynamic binaural experiments without a KEMAR or an anechoic robot facility.","The beamforming result indicates that algorithm evaluation must include motion as a condition, since even slow rotation changes performance in ways stationary tests miss."],"supporting_citations":[{"why":"Supplies the KEMAR HRTF measurements used as the acoustic realism reference for the printed head.","marker":"[5]"},{"why":"Shows that a retail mannequin provides reasonable acoustic shadowing, motivating a printable head design.","marker":"[24]"},{"why":"Proposes the earlier 3D-printed acoustic head simulator that this work refines.","marker":"[25]"},{"why":"Provides the bespoke dummy head whose mouth simulator design is adapted for the printed head.","marker":"[26]"},{"why":"Gives whispered speech spectra used as the benchmark showing motor noise stays at whisper level.","marker":"[35]"},{"why":"Supplies the RTF-steered binaural MVDR beamformer used in the dynamic beamforming application.","marker":"[36]"},{"why":"Defines the covariance whitening method used to estimate the relative transfer function in the beamformer.","marker":"[38]"},{"why":"Explains gear and motor harmonic noise structure that motivates the gear-free stepper design.","marker":"[20]"}],"fun_headline_variants":["Robotic dummy head brings repeatable motion to audio tests","Whisper-quiet robot head makes moving audio experiments repeatable","Low-cost robot head automates dynamic audio recordings","Robotic head on a quiet turntable repeats moving-source audio tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's motor-noise evidence comes from a microphone 1.0 m away, so the load-bearing assumption is that the ear microphones, mounted close to the turntable, also record clean audio while the head is moving.","fun_headline_variants_meta":{"raw":{"variants":["Robotic dummy head brings repeatable motion to audio tests","Whisper-quiet robot head makes moving audio experiments repeatable","Low-cost robot head automates dynamic audio recordings","Robotic head on a quiet turntable repeats moving-source audio tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1434,"prompt_tokens":837,"completion_tokens":597,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":453,"completion_tokens_details":{"reasoning_tokens":527}},"tokens_in":453,"tokens_out":597,"duration_ms":5485,"temperature":1.0,"reasoning_tokens":527,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:24:27.609170+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place a measurement microphone at or inside the ear canal of the rotating head in a quiet room and compare spectra with the motor still and rotating at 0.2 to 0.4 rev/s; if motor harmonics exceed the stationary noise floor at speech frequencies, the claim that recordings are uncontaminated and repeatable fails for those speeds.","supporting_citations":[{"cited_title":"HRFT Measurements of a KEMAR Dummy-head Microphone,","cited_arxiv_id":null,"evidence_quote":"Supplies the KEMAR HRTF measurements used as the acoustic realism reference for the printed head."},{"cited_title":"Acoustic impulse responses for wearable audio devices,","cited_arxiv_id":null,"evidence_quote":"Shows that a retail mannequin provides reasonable acoustic shadowing, motivating a printable head design."},{"cited_title":"3d-printed acoustic head simulators that talk and move,","cited_arxiv_id":null,"evidence_quote":"Proposes the earlier 3D-printed acoustic head simulator that this work refines."},{"cited_title":"Comparison of the acoustic effects of face masks on speech,","cited_arxiv_id":null,"evidence_quote":"Provides the bespoke dummy head whose mouth simulator design is adapted for the printed head."},{"cited_title":"A comparison of spectra of loud and whispered speech,","cited_arxiv_id":null,"evidence_quote":"Gives whispered speech spectra used as the benchmark showing motor noise stays at whisper level."},{"cited_title":"RTF-steered binaural MVDR beamforming incorporating an external microphone for dynamic acoustic scenarios,","cited_arxiv_id":null,"evidence_quote":"Supplies the RTF-steered binaural MVDR beamformer used in the dynamic beamforming application."},{"cited_title":"Performance analysis of the covariance subtraction method for relative transfer function estimation and comparison to the covariance whitening method,","cited_arxiv_id":null,"evidence_quote":"Defines the covariance whitening method used to estimate the relative transfer function in the beamformer."},{"cited_title":"Gear noise and the sideband phenomenon,","cited_arxiv_id":null,"evidence_quote":"Explains gear and motor harmonic noise structure that motivates the gear-free stepper design."}],"review_version":1}