{"id":"36dff593-d29e-422b-9335-a9aa382f9cc5","arxiv_id":"2506.02715","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A system uses AM-modulated ultrasound from off-the-shelf speakers and earphone DSP demodulation to deliver multi-channel, location-specific audio with preserved spatial cues, though no experimental validation is provided.","lead":"UltrasonicSpheres sends inaudible high-pitch audio from ordinary speakers to earphones that convert it back to sound, letting each user hear personalized content such as a museum narration in their chosen language while hearing their surroundings. The paper is a demo proposal that claims localized, multi-channel audio without special hardware, but it includes no measurements or user tests to back those claims.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Even ignoring the total absence of measurements, the multi-channel scheme as specified is internally inconsistent: with 4 kHz audio, AM sidebands of the 34 and 36 kHz carriers in Fig. 1 overlap, so the claimed simultaneous streams cannot be separated by the described band-pass filters.","rationale":"The reader's weakest assumption was the lack of hardware validation, specifically SPL/SNR figures for the off-the-shelf tweeter and earphone microphone chain. I partially agree with that concern, but the most load-bearing issue is sharper and internal: the multi-channel claim is geometrically impossible under the stated parameters. This is not merely an unverified empirical claim; it is an inconsistency between the paper's own Equation (1), the stated 4 kHz audio bandwidth, and the carrier frequencies shown in Figure 1. The overlap of the 34 kHz and 36 kHz channels means that no band-pass filter can isolate one stream from the other without audible crosstalk. Because this directly undercuts the abstract's central promise of multi-channel, location-specific audio delivery, it strengthens the rejection even if future hardware measurements were supplied. I nevertheless recommend leaving the reader's REJECT verdict unchanged: the paper also lacks any evaluation, and the channel-overlap flaw makes rejection robust.","tokens_in":5941,"tokens_out":7040,"duration_ms":86980,"concrete_test":"Run the paper's own modulation chain in simulation using Equation (1): take two speech-band signals (0-4 kHz), AM-modulate them onto 34 and 36 kHz carriers, sum the composite signal, then apply the described Chebyshev Type II band-pass filter centered at 34 kHz and the absolute-value/low-pass demodulator. Measure the output when only the 36 kHz channel is transmitting; if the reconstructed audio exceeds a standard intelligibility threshold (e.g., -20 dB crosstalk), the channel separation claim is refuted. The same test at carrier spacings of 8 and 10 kHz would confirm the required spacing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that multiple audio streams, e.g., English and German narrations, can be broadcast simultaneously from one speaker on distinct ultrasonic carriers and selected by the earphone's band-pass filter. The paper specifies in Section 3.1 that audio is low-pass filtered to 4 kHz (human speech) before AM, giving each carrier a passband of [fc-4, fc+4] kHz. Section 3.3 implements channel selection with a Chebyshev Type II band-pass filter followed by envelope demodulation. However, Figure 1 lists carriers at 26, 34, 36, and 44 kHz. The 34 kHz channel occupies 30-38 kHz and the 36 kHz channel occupies 32-40 kHz, overlapping over 32-38 kHz. A band-pass filter centered at 34 kHz will pass part of the 36 kHz channel, and vice versa; after rectification and low-pass filtering this produces audible crosstalk. Even the 30 kHz carrier chosen in Equation (1) overlaps the 34 kHz channel (26-34 vs. 30-38 kHz). Thus the multi-channel portion of the central claim fails at the level of the paper's own design, independent of whether any hardware prototype was measured. A correct design would require carrier separation greater than 8 kHz, not the 2 kHz shown in the figure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"UltrasonicSpheres proposes a system for location-specific audio delivery in which audio signals are AM-modulated onto ultrasonic carriers, broadcast from off-the-shelf tweeters, and demodulated by OpenEarable 2.0 earphones using band-pass filtering and envelope detection. The paper claims that multiple simultaneous streams (e.g., English and German narrations) can be selected by carrier frequency, that spatial audio perception is preserved through interaural time differences, and that the earphones remain acoustically transparent. It describes a museum-like demo with two exhibits and two language channels, but reports no measurements or user evaluations.","tokens_in":6256,"tokens_out":4803,"duration_ms":48438,"significance":"If validated, the approach would offer a low-cost, accessible alternative to parametric-array and metasurface systems for localized audio, using commodity tweeters and an open-source earable platform. The paper's strengths are its clear system concept, the use of standard AM theory, and the concrete identification of a receiver platform (OpenEarable 2.0) that could make the idea reproducible. However, none of the central experiential claims are currently supported by measured data, and the multi-channel design as specified has a carrier-spacing inconsistency.","major_comments":[{"comment":"The multi-channel separation is internally inconsistent. With the 4 kHz speech low-pass filter, an AM channel at carrier f_c occupies [f_c−4, f_c+4] kHz. The 34 kHz and 36 kHz channels shown in Figure 1 therefore overlap over 32–38 kHz, and the 30 kHz carrier used in Eq. (1) overlaps the 34 kHz channel over 30–34 kHz. A Chebyshev Type II band-pass filter cannot separate these channels, so the claimed simultaneous English/German streams would produce audible crosstalk after envelope detection. The design needs carrier separations of at least 8 kHz (with practical margin, larger), a narrower audio bandwidth, or a different multiplexing scheme, and the figure and equations must use a consistent carrier set.","section":"Section 3.1 and Figure 1"},{"comment":"No quantitative or perceptual evaluation supports the central claims. The paper reports no sound-pressure-level measurements of the ultrasonic emission, no SNR at the earphone microphone, no measurement of demodulated audio intelligibility or distortion, and no user study of localization, channel selection, or acoustic transparency. Without these, the Abstract's claims that users \"can demodulate their selected stream\" and that the system \"preserves spatial audio perception\" are unsupported. At minimum, the paper should include measured tweeter and microphone frequency responses, received SNR as a function of distance and angle, demodulated audio spectrograms or sample excerpts, and a small listening test.","section":"Sections 3 and 4"},{"comment":"The demodulation description is internally inconsistent and incomplete. Eq. (3) applies a band-pass filter to |y(t)|, whereas the text and Figure 2 say the absolute value is followed by a low-pass filter to remove high-frequency components. Since |cos(2π f_c t)| contains harmonics at multiples of 2f_c, the filter type and cutoff must be specified. The paper should also analyze the distortion introduced by envelope detection, including the condition 1+k_a x(t) > 0 and the effect of nearby carriers, or report measured total harmonic distortion.","section":"Section 3.2, Eq. (3)"},{"comment":"The claim that spatial perception is preserved because interaural time difference is preserved is not established. The receiver path includes per-ear filtering, envelope detection, and dynamic range compression, all of which modify interaural time and level differences; no latency, phase, or localization test is reported. Similarly, the claim that the earphones are \"acoustically transparent\" needs verification, for example through occlusion-effect or insertion-gain measurements.","section":"Sections 1 and 3.2"}],"minor_comments":[{"comment":"The DSP parameters are not given: bell-curve equalizer settings, Chebyshev filter order and cutoffs, limiter thresholds, and the amplification factor A in Eq. (3). A table of these parameters would greatly improve reproducibility.","section":"Section 3.3"},{"comment":"Eq. (1) uses a 30 kHz carrier, but Figure 1 lists carriers at 26, 34, 36, and 44 kHz. The paper should state which carrier set the prototype actually uses and why the text and figure differ.","section":"Section 3.1 and Figure 1"},{"comment":"The label \"26 kHZ\" should be \"26 kHz,\" and the annotation \"Font Size\" appears to be a stray UI element rather than a meaningful caption component.","section":"Figure 1"},{"comment":"The demo description lacks concrete geometry and safety information: speaker-to-exhibit distances, room dimensions, playback levels, and whether ultrasonic exposure remains within relevant safety guidelines.","section":"Section 4"},{"comment":"The header \"Unpublished working draft. Not for distribution.\" appears on the manuscript; if this is a status marker it should be removed before submission, or clarified in the submission metadata.","section":"Title page"}],"recommendation":"major_revision","confidential_remarks":"There is no measurement or user study in the paper, and the multi-channel carrier-spacing issue is a concrete design error, not merely a missing evaluation. Both are fixable within the manuscript's scope: the carrier set can be re-spaced, and basic acoustic measurements plus a small listening test would substantially strengthen the claims. The editor may also wish to confirm the intended submission status given the \"unpublished working draft\" header. The overlap of the receiver platform with the authors' own OpenEarable 2.0 is not itself a problem, but the paper should provide enough independent detail (filter parameters, speaker model, amplifier settings) to make the demo reproducible by others."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the good news: the paper is honest about its scope. It's a demo paper, and it reads like one. The idea is to amplitude-modulate speech onto a 30 kHz carrier, broadcast it from a bookshelf tweeter, and pick it up with an open-ear microphone on the OpenEarable platform, demodulating by absolute value in the DSP. The basic receive trick was shown in UltrasonicWhisper (which they cite), so the new bit is the multi-channel carrier scheme and the claim that ordinary tweeters can transmit.\n\nWhat the paper does well: it lays out the signal chain clearly, from LPF to AM to BPF to envelope detection, and it doesn't oversell beyond the demo context. The argument that interaural time difference is preserved across the ultrasonic link is plausible. Related work is fine, and self-citing OpenEarable is appropriate since the platform is the actual receiver.\n\nSoft spots: first, the multi-channel plan has a concrete design error. With 4 kHz speech, AM sidebands occupy fc±4 kHz. The carriers in Figure 1 are 26, 34, 36, and 44 kHz. The 34 kHz channel occupies 30–38 kHz and the 36 kHz channel occupies 32–40 kHz; they overlap over 32–38 kHz. The specified Chebyshev band-pass cannot separate them, so the two languages would bleed into each other. The 26 and 34 channels just touch at 30 kHz with zero guard band. This is independent of hardware quality; the design needs carrier spacing over 8 kHz (or narrower audio bandwidth) to work.\n\nSecond, there is no evaluation at all. No SPL measurements of the tweeter at 30 kHz, no SNR or distortion figures at the mic, no intelligibility or localization experiment. The phrase 'we found that many consumer audio tweeters can serve this purpose' is the only evidence, and it appears without numbers. For a demo paper this might be tolerable if the demo is the artifact, but for a scientific preprint it leaves the central claims unsupported.\n\nThird, the envelope detector via absolute value is fine when the carrier is dominant, but with overlapping sidebands it will demodulate the interfering channel too. A synchronous detector with carrier recovery would be more selective, but they don't use one.\n\nOverall, this is a plausible idea with a real bug in the multi-channel scheme and no measurements. It should not be rejected for lack of novelty—the combination is a legitimate extension—but it needs either a live demo with audibility tests or a corrected design and some basic measurements before it supports the claims.\n\nWho is this for? People building earable audio systems or considering ultrasonic data channels in public spaces. The related-work section is a useful short survey.\n\nRecommendation: I would not send this to a full archival venue as is. If the venue has a demo track where the artifact is the submission, that's the right home, but as a peer-reviewed paper it needs a significant revision. I'd desk-reject rather than burn referee time on a paper with a self-inconsistent multi-channel design and zero data.","headline":"A clever but unvalidated demo concept whose multi-channel scheme as specified cannot work with the chosen carrier frequencies.","tokens_in":6756,"tokens_out":3513,"would_cite":false,"duration_ms":36544,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"UltrasonicSpheres shows that ordinary tweeters can broadcast multi-channel ultrasound that earphone DSPs demodulate into localized, selectable audio.","keywords":["ultrasound","earables","sound zones","spatial audio","amplitude modulation","museum audio","public spaces","hearables"],"falsifier":"Place an unmodified bookshelf tweeter in a typical room, play speech amplitude-modulated onto a 30 kHz carrier, and record the demodulated output from the earphone DSP while varying distance and angle; if the received carrier level drops to the ambient noise floor or the recovered speech is unintelligible at the intended sphere radius, the central premise is refuted.","tokens_in":5784,"feed_emoji":"🎧","tokens_out":10244,"duration_ms":97015,"temperature":0.7,"pith_summary":"UltrasonicSpheres aims to deliver location-specific audio to a listener's earphones using equipment that is already common: a bookshelf tweeter and an earable with a microphone and DSP. Audio is amplitude-modulated onto a 30 kHz ultrasonic carrier; the earphone's DSP isolates the carrier, rectifies it, and low-pass filters it to recover the audio, so the listener hears what the speaker 'says' without the ultrasonic signal being audible to the naked ear. Because each ear is demodulated independently, the interaural time difference is preserved and the sound seems to come from the physical speaker. Multiple channels (for example, German and English narrations) can be broadcast at once on different carrier frequencies and selected on the earable. If the system works as described, museum visitors could walk through exhibits and hear their chosen narration while staying fully aware of ambient sound.","feed_headline":"Earphones decode off-the-shelf tweeters into personal sound spheres","feed_subtitle":"Multilingual narration rides separate ultrasonic channels while ambient sound and spatial cues stay intact.","key_machinery":"The carrying mechanism is amplitude modulation with a 30 kHz ultrasonic carrier followed by envelope detection in the earphone. The modulated signal is $y(t)=[1+k_a x(t)]\\cos(2\\pi f_c t)$; after a band-pass filter selects the desired channel from several carriers, the absolute value $|y(t)|$ rectifies the waveform, a low-pass filter removes the residual $2f_c$ components, and amplification plus limiting yields the recovered audio. This envelope-detection path avoids phase-locked carrier recovery, which keeps the receiver simple. The DSP pipeline also includes an equalizer stage to compensate for the tweeter's ultrasonic frequency response, and the recovered audio is mixed back with ambient sound so the earphone stays acoustically transparent.","core_discovery":"The paper's central claim is that an off-the-shelf high-fidelity tweeter, driven by a 96 kHz audio interface, radiates enough ultrasonic energy for an earphone-mounted microphone and DSP to demodulate into intelligible audio. The transmitter sends $y(t)=[1+k_a x(t)]\\cos(2\\pi f_c t)$ with carrier $f_c=30$ kHz, keeping the lower sideband at 26 kHz, just above human hearing. The receiver band-pass filters one carrier from a composite of several, takes the absolute value to detect the envelope, low-pass filters to recover the baseband audio, and mixes it with ambient sound before playback. The paper argues that because the signal travels at the speed of sound and is demodulated separately in each ear, the listener's interaural time difference is preserved, so the audio is perceived as originating at the speaker's physical location. The demo uses two speakers, one broadcasting a soundscape and one broadcasting English and German narrations on separate carriers, to show co-located, personalized, spatially anchored audio.","pith_inferences":["The same carrier-multiplexed envelope detection could be tested beyond narration, for example directional alerts or adaptive soundscapes that respond to head orientation, since per-ear demodulation preserves interaural cues.","A quantitative speech-intelligibility measurement across distance, angle, and ambient-noise level would tell whether the approach generalizes from a demo to real exhibition spaces; the paper reports no such measurements.","Because earphones demodulate ultrasound, the system also sits on a security surface: a tweeter in a public space could inject audio into a listener's ear, and the channel-separation and filtering stages here suggest where countermeasures would live.","The number of simultaneous channels is bounded by the tweeter's ultrasonic bandwidth and the microphone's 80 kHz ceiling, so a concrete channel-capacity estimate is a natural next step."],"forward_implications":["Visitors not wearing the earphones are unaffected, because the 30 kHz carrier and its sidebands are inaudible to human ears.","A listener can switch between co-located streams, such as German and English narration, by tapping the earphones, without pairing, tracking, or extra infrastructure.","Because each ear demodulates independently, the interaural time difference survives, so the sound is perceived as coming from the physical speaker rather than from inside the earphone.","Ambient sound is mixed back into the earpiece, so users remain spatially and situationally aware while receiving personalized audio.","The system needs only a commodity tweeter and an open-source earable platform, which suggests it could be deployed far more cheaply than parametric arrays or phased-array ultrasound systems."],"supporting_citations":[{"why":"Supplies the open-source earphone platform, DSP, and ultrasonic-capable microphone that the receiver prototype is built on.","marker":"[10]"},{"why":"Demonstrates that commodity earphones demodulate ultrasound into audible sound, the capability UltrasonicSpheres repurposes for curated audio delivery.","marker":"[12]"},{"why":"Represents the commercial parametric-speaker baseline the paper contrasts with its cheaper, off-the-shelf approach.","marker":"[6]"},{"why":"Reviews parametric-array demodulation, the alternative technique that requires specialized, high-intensity ultrasonic emitters the paper avoids.","marker":"[8]"}],"fun_headline_variants":["Off-the-shelf tweeters send spatial audio only earables can hear","Earphone mic decodes ultrasonic tweeters into private sound zones","Ultrasonic spheres: multi-channel audio with zero pairing or tracking","Localized sound from tweeters, decoded by earable microphones","Inaudible to others, audible to you: ultrasonic tweeter audio"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system depends on an unmodified bookshelf tweeter radiating enough 30 kHz energy that the earphone microphone can demodulate it into intelligible audio at listening distances; the paper asserts this capability but reports no sound-pressure-level, signal-to-noise, or audio-quality measurements.","fun_headline_variants_meta":{"raw":{"variants":["Off-the-shelf tweeters send spatial audio only earables can hear","Earphone mic decodes ultrasonic tweeters into private sound zones","Ultrasonic spheres: multi-channel audio with zero pairing or tracking","Localized sound from tweeters, decoded by earable microphones","Inaudible to others, audible to you: ultrasonic tweeter audio"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000291,"raw_usage":{"total_tokens":1690,"prompt_tokens":928,"completion_tokens":762,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":669}},"tokens_in":544,"tokens_out":762,"duration_ms":6846,"temperature":1.0,"reasoning_tokens":669,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:17:04.082796+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place an unmodified bookshelf tweeter in a typical room, play speech amplitude-modulated onto a 30 kHz carrier, and record the demodulated output from the earphone DSP while varying distance and angle; if the received carrier level drops to the ambient noise floor or the recovered speech is unintelligible at the intended sphere radius, the central premise is refuted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the open-source earphone platform, DSP, and ultrasonic-capable microphone that the receiver prototype is built on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Demonstrates that commodity earphones demodulate ultrasound into audible sound, the capability UltrasonicSpheres repurposes for curated audio delivery."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the commercial parametric-speaker baseline the paper contrasts with its cheaper, off-the-shelf approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reviews parametric-array demodulation, the alternative technique that requires specialized, high-intensity ultrasonic emitters the paper avoids."}],"review_version":1}