Pith. sign in

REVIEW 4 major objections 5 minor 15 references

UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables

T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read UltrasonicSpheres shows that ordinary tweeters can broadcast multi-channel ultrasound that earphone DSPs demodulate into localized, selectable audio.

desk verdict A clever but unvalidated demo concept whose multi-channel scheme as specified cannot work with the chosen carrier frequencies. read the letter →

arxiv 2506.02715 v2 pith:VUMZJV5U submitted 2025-06-03 cs.SD cs.HCeess.AS

classification cs.SDcs.HCeess.AS
keywords ultrasoundearablessoundzonesspatialaudioamplitudemodulationmuseumpublicspaceshearables
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

UltrasonicSpheres aims to deliver location-specific audio to a listener's earphones using equipment that is already common: a bookshelf tweeter and an earable with a microphone and DSP. Audio is amplitude-modulated onto a 30 kHz ultrasonic carrier; the earphone's DSP isolates the carrier, rectifies it, and low-pass filters it to recover the audio, so the listener hears what the speaker 'says' without the ultrasonic signal being audible to the naked ear. Because each ear is demodulated independently, the interaural time difference is preserved and the sound seems to come from the physical speaker. Multiple channels (for example, German and English narrations) can be broadcast at once on different carrier frequencies and selected on the earable. If the system works as described, museum visitors could walk through exhibits and hear their chosen narration while staying fully aware of ambient sound.

What carries the argument

The carrying mechanism is amplitude modulation with a 30 kHz ultrasonic carrier followed by envelope detection in the earphone. The modulated signal is $y(t)=[1+k_a x(t)]\cos(2\pi f_c t)$; after a band-pass filter selects the desired channel from several carriers, the absolute value $|y(t)|$ rectifies the waveform, a low-pass filter removes the residual $2f_c$ components, and amplification plus limiting yields the recovered audio. This envelope-detection path avoids phase-locked carrier recovery, which keeps the receiver simple. The DSP pipeline also includes an equalizer stage to compensate for the tweeter's ultrasonic frequency response, and the recovered audio is mixed back with ambient sound so the earphone stays acoustically transparent.

What would settle it

Place an unmodified bookshelf tweeter in a typical room, play speech amplitude-modulated onto a 30 kHz carrier, and record the demodulated output from the earphone DSP while varying distance and angle; if the received carrier level drops to the ambient noise floor or the recovered speech is unintelligible at the intended sphere radius, the central premise is refuted.

Watch

Extended reading notes

Core claim

The paper's central claim is that an off-the-shelf high-fidelity tweeter, driven by a 96 kHz audio interface, radiates enough ultrasonic energy for an earphone-mounted microphone and DSP to demodulate into intelligible audio. The transmitter sends $y(t)=[1+k_a x(t)]\cos(2\pi f_c t)$ with carrier $f_c=30$ kHz, keeping the lower sideband at 26 kHz, just above human hearing. The receiver band-pass filters one carrier from a composite of several, takes the absolute value to detect the envelope, low-pass filters to recover the baseband audio, and mixes it with ambient sound before playback. The paper argues that because the signal travels at the speed of sound and is demodulated separately in each ear, the listener's interaural time difference is preserved, so the audio is perceived as originating at the speaker's physical location. The demo uses two speakers, one broadcasting a soundscape and one broadcasting English and German narrations on separate carriers, to show co-located, personalized, spatially anchored audio.

Load-bearing premise

The system depends on an unmodified bookshelf tweeter radiating enough 30 kHz energy that the earphone microphone can demodulate it into intelligible audio at listening distances; the paper asserts this capability but reports no sound-pressure-level, signal-to-noise, or audio-quality measurements.

Editorial extensions

If this is right

  • Visitors not wearing the earphones are unaffected, because the 30 kHz carrier and its sidebands are inaudible to human ears.
  • A listener can switch between co-located streams, such as German and English narration, by tapping the earphones, without pairing, tracking, or extra infrastructure.
  • Because each ear demodulates independently, the interaural time difference survives, so the sound is perceived as coming from the physical speaker rather than from inside the earphone.
  • Ambient sound is mixed back into the earpiece, so users remain spatially and situationally aware while receiving personalized audio.
  • The system needs only a commodity tweeter and an open-source earable platform, which suggests it could be deployed far more cheaply than parametric arrays or phased-array ultrasound systems.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same carrier-multiplexed envelope detection could be tested beyond narration, for example directional alerts or adaptive soundscapes that respond to head orientation, since per-ear demodulation preserves interaural cues.
  • A quantitative speech-intelligibility measurement across distance, angle, and ambient-noise level would tell whether the approach generalizes from a demo to real exhibition spaces; the paper reports no such measurements.
  • Because earphones demodulate ultrasound, the system also sits on a security surface: a tweeter in a public space could inject audio into a listener's ear, and the channel-separation and filtering stages here suggest where countermeasures would live.
  • The number of simultaneous channels is bounded by the tweeter's ultrasonic bandwidth and the microphone's 80 kHz ceiling, so a concrete channel-capacity estimate is a natural next step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. UltrasonicSpheres proposes a system for location-specific audio delivery in which audio signals are AM-modulated onto ultrasonic carriers, broadcast from off-the-shelf tweeters, and demodulated by OpenEarable 2.0 earphones using band-pass filtering and envelope detection. The paper claims that multiple simultaneous streams (e.g., English and German narrations) can be selected by carrier frequency, that spatial audio perception is preserved through interaural time differences, and that the earphones remain acoustically transparent. It describes a museum-like demo with two exhibits and two language channels, but reports no measurements or user evaluations.

Significance. If validated, the approach would offer a low-cost, accessible alternative to parametric-array and metasurface systems for localized audio, using commodity tweeters and an open-source earable platform. The paper's strengths are its clear system concept, the use of standard AM theory, and the concrete identification of a receiver platform (OpenEarable 2.0) that could make the idea reproducible. However, none of the central experiential claims are currently supported by measured data, and the multi-channel design as specified has a carrier-spacing inconsistency.

major comments (4)
  1. [Section 3.1 and Figure 1] The multi-channel separation is internally inconsistent. With the 4 kHz speech low-pass filter, an AM channel at carrier f_c occupies [f_c−4, f_c+4] kHz. The 34 kHz and 36 kHz channels shown in Figure 1 therefore overlap over 32–38 kHz, and the 30 kHz carrier used in Eq. (1) overlaps the 34 kHz channel over 30–34 kHz. A Chebyshev Type II band-pass filter cannot separate these channels, so the claimed simultaneous English/German streams would produce audible crosstalk after envelope detection. The design needs carrier separations of at least 8 kHz (with practical margin, larger), a narrower audio bandwidth, or a different multiplexing scheme, and the figure and equations must use a consistent carrier set.
  2. [Sections 3 and 4] No quantitative or perceptual evaluation supports the central claims. The paper reports no sound-pressure-level measurements of the ultrasonic emission, no SNR at the earphone microphone, no measurement of demodulated audio intelligibility or distortion, and no user study of localization, channel selection, or acoustic transparency. Without these, the Abstract's claims that users "can demodulate their selected stream" and that the system "preserves spatial audio perception" are unsupported. At minimum, the paper should include measured tweeter and microphone frequency responses, received SNR as a function of distance and angle, demodulated audio spectrograms or sample excerpts, and a small listening test.
  3. [Section 3.2, Eq. (3)] The demodulation description is internally inconsistent and incomplete. Eq. (3) applies a band-pass filter to |y(t)|, whereas the text and Figure 2 say the absolute value is followed by a low-pass filter to remove high-frequency components. Since |cos(2π f_c t)| contains harmonics at multiples of 2f_c, the filter type and cutoff must be specified. The paper should also analyze the distortion introduced by envelope detection, including the condition 1+k_a x(t) > 0 and the effect of nearby carriers, or report measured total harmonic distortion.
  4. [Sections 1 and 3.2] The claim that spatial perception is preserved because interaural time difference is preserved is not established. The receiver path includes per-ear filtering, envelope detection, and dynamic range compression, all of which modify interaural time and level differences; no latency, phase, or localization test is reported. Similarly, the claim that the earphones are "acoustically transparent" needs verification, for example through occlusion-effect or insertion-gain measurements.
minor comments (5)
  1. [Section 3.3] The DSP parameters are not given: bell-curve equalizer settings, Chebyshev filter order and cutoffs, limiter thresholds, and the amplification factor A in Eq. (3). A table of these parameters would greatly improve reproducibility.
  2. [Section 3.1 and Figure 1] Eq. (1) uses a 30 kHz carrier, but Figure 1 lists carriers at 26, 34, 36, and 44 kHz. The paper should state which carrier set the prototype actually uses and why the text and figure differ.
  3. [Figure 1] The label "26 kHZ" should be "26 kHz," and the annotation "Font Size" appears to be a stray UI element rather than a meaningful caption component.
  4. [Section 4] The demo description lacks concrete geometry and safety information: speaker-to-exhibit distances, room dimensions, playback levels, and whether ultrasonic exposure remains within relevant safety guidelines.
  5. [Title page] The header "Unpublished working draft. Not for distribution." appears on the manuscript; if this is a status marker it should be removed before submission, or clarified in the submission metadata.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the system uses standard AM/envelope-demodulation equations, and the OpenEarable self-citation is a hardware platform reference, not a load-bearing mathematical input.

full rationale

UltrasonicSpheres does not derive any claimed result from a fitted parameter or from a self-citation chain. The modulation and demodulation equations in Section 3 are standard amplitude-modulation and envelope-detection expressions; Eq. (3) is an envelope-detector approximation, not a model fitted to the paper's own output. The choice of a 30 kHz carrier and a 4 kHz audio low-pass filter is a stated design decision with a bandwidth rationale, not a parameter fit disguised as a prediction. Multi-channel operation is described as summing individually modulated signals and isolating them with band-pass filters; no quantity is fitted to a subset of data and then relabeled as a predicted outcome. The only self-citation is Reference [10], OpenEarable 2.0, used as the receiver hardware platform. That platform is open-source and is cited as an existing implementation resource rather than as a theorem or uniqueness argument, so the citation does not make the central claim circular. The apparent carrier-overlap problem (e.g., 34 kHz and 36 kHz carriers with 4 kHz audio) is an internal-consistency or correctness concern, not a circularity, because it does not involve the paper's conclusion being equivalent to its premises by construction. Accordingly, no circular steps are identified and the circularity score is 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The system description relies on several unverified domain assumptions about consumer tweeter ultrasonic output, earphone microphone sensitivity, envelope-detection fidelity, and interaural cue preservation, plus multiple unspecified DSP gain and filter parameters. No measurements, simulations, or user studies are included.

free parameters (5)
  • Carrier frequency f_c = 30 kHz
    Chosen so the lower sideband (26 kHz) remains above the audible range; not fitted to data.
  • Speech low-pass cutoff = 4 kHz
    Set to cover human speech spectrum; arbitrary upper bound for content.
  • Demodulation gain A = not specified
    Custom amplification factor in Eq. (3) to scale the recovered audio; value is not given and would need tuning.
  • Bell-curve equalizer parameters = not specified
    DSP equalizer used to compensate ultrasonic frequency response; no values provided.
  • Limiter thresholds = not specified
    Dynamic range compression parameter to suppress clicks; no values provided.
assumptions (5)
  • domain assumption Standard tweeters can output usable energy at 30 kHz with sufficient SPL after propagation.
    Invoked in Section 3 (Implementation) when selecting off-the-shelf speakers; no measurements provided.
  • domain assumption The earphone microphone (SPH0641LU4H-1) and ADC at 192 kHz capture the ultrasonic carrier with adequate SNR.
    Specified in Section 3.3, but no sensitivity data at 30 kHz is given.
  • domain assumption Envelope detection via absolute value faithfully recovers the audio signal for intelligible listening.
    Used in Eq. (3) without analysis of harmonic distortion or intermodulation from multiple carriers.
  • domain assumption Interaural time differences of the modulated envelope are preserved, so users localize the sound at the speaker.
    Assumed in Section 3, paragraph 2; no listening test or localization measurement is reported.
  • domain assumption Multiple ultrasonic channels on different carriers do not interfere.
    Implied by the multi-channel claim in Section 3 and Demo, but no crosstalk measurements are given.

how reviews work

0 comments
Cite this review

Pith. "Pith review of UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables." pith.science (2026). https://pith.science/paper/VUMZJV5U

@misc{pith2026250602715,
  author       = {Pith},
  title        = {Pith review of: UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VUMZJV5U}},
  note         = {Machine review of arXiv:2506.02715}
}
read the original abstract

We present a demo of UltrasonicSpheres, a novel system for location-specific audio delivery using wearable earphones that decode ultrasonic signals into audible sound. Unlike conventional beamforming setups, UltrasonicSpheres relies on single ultrasonic speakers to broadcast localized audio with multiple channels, each encoded on a distinct ultrasonic carrier frequency. Users wearing our acoustically transparent earphones can demodulate their selected stream, such as exhibit narrations in a chosen language, while remaining fully aware of ambient environmental sounds. The experience preserves spatial audio perception, giving the impression that the sound originates directly from the physical location of the source. This enables personalized, localized audio without requiring pairing, tracking, or additional infrastructure. Importantly, visitors not equipped with the earphones are unaffected, as the ultrasonic signals are inaudible to the human ear. Our demo invites participants to explore multiple co-located audio zones and experience how UltrasonicSpheres supports unobtrusive delivery of personalized sound in public spaces.

Figures

Figures reproduced from arXiv: 2506.02715 by the authors.

Figure 2
Figure 2. The original signal is modulated and transmitted via a speaker, where it becomes mixed with environmental sound. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of the magnitude spectrum in the fre [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figure 5
Figure 5. (Left) A visitor experiencing a localized Ultrasonic [PITH_FULL_IMAGE:figures/full_fig_p004_5.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

15 extracted references · 15 canonical work pages

  1. [1]

    Apple Inc. 2025. Control Spatial Audio and head tracking. https://support.apple. com/en-en/guide/airpods/dev00eb7e0a3/web Accessed: 2025-05-24

  2. [2]

    Terence Betlehem, Wen Zhang, Mark A Poletti, and Thushara D Abhayapala

  3. [3]

    Tuochao Chen, Malek Itani, Sefik Emre Eskimez, Takuya Yoshioka, and Shyam- nath Gollakota. 2024. Hearable devices with sound bubbles. Nature Electronics (2024), 1–12

  4. [4]

    Xiaoran Fan, David Pearl, Richard Howard, Longfei Shangguan, and Trausti Thormundsson. 2023. Apg: Audioplethysmography for cardiac monitoring in hearables. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking. 1–15

  5. [5]

    Yang Gao, Wei Wang, Vir V Phoha, Wei Sun, and Zhanpeng Jin. 2019. EarEcho: Using ear canal echo for wearable authentication. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 3, 3 (2019), 1–24

  6. [6]

    Holosonics. 2025. Audio Spotlight by Holosonics. https://www.holosonics.com/. Accessed: 2025-05-22. Unpublished working draft.Not for distribution. UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables

  7. [7]

    Yincheng Jin, Yang Gao, Yanjun Zhu, Wei Wang, Jiyang Li, Seokmin Choi, Zhangyu Li, Jagmohan Chauhan, Anind K Dey, and Zhanpeng Jin. 2021. Son- icASL: An acoustic-based sign language gesture recognizer using earphones. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technolo- gies 5, 2 (2021), 1–30

  8. [8]

    Shangming Mei, Hui Xu, Yihua Hu, Mohammed Alkahtani, and Yangang Wang

Show all 15 references
  1. [9]

    Yoichi Ochiai, Takayuki Hoshi, and Ippei Suzuki. 2017. Holographic whisper: Rendering audible sound spots in three-dimensional space by focusing ultrasonic waves. In proceedings of the 2017 CHI conference on human factors in computing systems. 4314–4325

  2. [10]

    Tobias Röddiger, Michael Küttner, Philipp Lepold, Tobias King, Dennis Moschina, Oliver Bagge, Joseph A Paradiso, Christopher Clarke, and Michael Beigl. 2025. OpenEarable 2.0: Open-Source Earphone Platform for Physiological Ear Sens- ing. Proceedings of the ACM on Interactive, ...

  3. [11]

    Bandhav Veluri, Malek Itani, Tuochao Chen, Takuya Yoshioka, and Shyamnath Gollakota. 2024. Look once to hear: Target speech hearing with noisy examples. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 1–16

  4. [12]

    Hiroki Watanabe and Tsutomu Terada. 2023. UltrasonicWhisper: Ultrasound Can Generate Audible Sound in Your Hearable. In Proceedings of the 2023 ACM International Symposium on Wearable Computers . 132–134

  5. [13]

    Jia-Xin Zhong, Jun Ji, Xiaoxing Xia, Hyeonu Heo, and Yun Jing. 2025. Audible enclaves crafted by nonlinear self-bending ultrasonic beams. Proceedings of the National Academy of Sciences 122, 12 (2025), e2408975122. Received 30 May 2025

  6. [2015]

    IEEE Signal Processing Magazine 32, 2 (2015), 81–91

    Personal sound zones: Delivering interface-free audio to multiple listeners. IEEE Signal Processing Magazine 32, 2 (2015), 81–91

  7. [2022]

    In International Joint Conference on Energy, Electrical and Power Engineering

    The Parametric Array Speaker: A Review. In International Joint Conference on Energy, Electrical and Power Engineering . Springer, 1254–1271

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.