{"id":"d8f9db74-aea0-430d-a649-e09a1eb7af44","arxiv_id":"2506.17570","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Electromagnetic emanations from a VR headset can be classified with a nearby software-defined radio and a fine-tuned ResNet to identify the running VR app and the user's activity, with a claimed accuracy near 99%.","lead":"This paper shows that a passive radio receiver placed one to two meters away can capture electromagnetic noise leaking from a VR headset and, with a machine-learning classifier, identify which of 15 virtual reality apps is running and which of 4 user activities is happening. The reported accuracy is about 99%, but the evaluation lacks error bars and cross-user tests, so that number is not yet trustworthy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 99% accuracy is built on random splits of same-session IQ chunks, with no cross-session or cross-user test reported, so the claimed app/activity signatures may be session artifacts rather than generalizable EM fingerprints.","rationale":"I read the paper as a feasibility claim: a passive EM sniffer can identify which of 15 VR apps is running and which of 4 activities the user is performing, with 99% accuracy. For that claim to hold, the learned features must be caused by the app or activity, not by the particular recording session. The most load-bearing weak point is the evaluation protocol. The text reports data collected with \"the VR user\" in two rooms, but gives no participant count, no session count, and no temporal separation between training and test chunks. Splitting 500K-sample chunks randomly from continuous IQ streams means train and test chunks are interleaved in time and share the same physical context, so the model can discriminate sessions or recording-order artifacts instead of apps. The paper's own Section 8.2 limits the attack to a single user, confirming that multi-user generalization was not addressed. The orientation microbenchmark strengthens the concern: at side orientations the measured USNR is below 1 dB, yet the classifier still reports about 0.96 accuracy; that is not credible purely from claimed emanation spikes and suggests the model is exploiting stable contextual cues. The reader's weakest_assumption is exactly this generalization and invariance problem, and I agree with it. I do not count the lack of formal proof as a problem; this is an empirical systems paper. The real SDR measurements, the two-headset comparison, the ML-model ablation, and the IRB-approved collection are genuine evidence. The issue is that the evidence is not structured to support the generalization claim. The miscalculated USNR gain in Section 5.1 is also a real flaw, but it is secondary: the central 99% claim collapses on the evaluation-protocol problem without needing the enhancement math. A leave-one-session-out test is the single check that would settle whether the signatures are stable. If that test passes, the rejection should be revisited; if it fails, the 99% number is an artifact. Because the reader already recommended REJECT and my concern is the same one, I recommend no change to the verdict.","tokens_in":21015,"tokens_out":4885,"duration_ms":59529,"concrete_test":"Require the authors to report participant and session counts, then rerun the evaluation with leave-one-session-out splits: train on all 500K-sample chunks from every session except one, then test only on the held-out session, recorded at a different time with the antenna repositioned and, if available, a different user and room. Compare that held-out accuracy with the 70/15/15 random-split accuracy in Section 7.1. If held-out accuracy falls materially below 99% or toward chance, the reported accuracy is session leakage and the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"VReaves's central claim is that passively sniffed EM spectra identify VR apps and activities at 99%. That claim requires the extracted FFT/STFT signatures to be stable across sessions, users, and rooms. The paper never tests this. In Section 6 (\"Experimental settings\") and Section 7.1 (\"Method\"), IQ samples are collected with \"the VR user\" in two rooms, divided into 500K-sample chunks, and randomly split 70/15/15 into train/validation/test. The 99% figures in Figs. 20 and 21 are from that random split. Chunks from a continuous recording share the same headset, user position, antenna placement, room RF background, and time period; a random split therefore leaks session context into the training set and can produce near-perfect accuracy even if the model has learned nothing about the app. No participant count, session count, or cross-session/cross-user experiment is reported, and Section 8.2 explicitly restricts the work to a single targeted user. Fig. 31 sharpens the problem: app identification and activity recognition stay near 0.96 at orientations where Fig. 30 reports USNR below 1 dB. If the emanation spikes are below the noise floor, the classifier must be relying on some other stable cue; the most parsimonious explanation is recording context rather than the app's computational signature. The central attack claim therefore rests on an untested generalization assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes VReaves, a passive electromagnetic (EM) side-channel attack against VR headsets. The authors use a USRP N210 and directional antenna to capture unintentional emanations from a Meta Quest 3 and HTC VIVE XR Elite, then apply a signal-processing pipeline of noise-floor smoothing, ambient interference subtraction, and averaging FFT/STFT to produce frequency-domain and spectrogram features. A fine-tuned ResNet is trained to identify 15 VR app identities and 4 app activities (entering, configuring, running, exiting), with reported accuracies around 99%. The paper also reports microbenchmarks for distance, orientation, number of frequency bands, emanation duration, and headset model, plus a case study with the receiver hidden in a backpack.","tokens_in":21282,"tokens_out":8518,"duration_ms":91366,"significance":"If the 99% accuracy held across users, sessions, and environments, VReaves would be a notable privacy threat: it would enable passive, malware-free identification of VR app identity and user activity from a distance using COTS hardware. The work is among the first to target VR headsets with EM side channels, and it has concrete strengths: real SDR measurements, two commercial headsets, a detailed signal-processing pipeline, a comparison against LSTM/transformer baselines, and an explicit limitations section (Sec. 8.2). I do not see a circularity problem in Eqs. (1)-(5); they are standard signal models, and the reported accuracy is an empirical supervised-learning result rather than a quantity derived from those equations. The central weakness is that the evaluation does not separate sessions or users, so the reported accuracy may reflect recording context rather than stable app/activity signatures.","major_comments":[{"comment":"The reported 99% accuracy is based on a single random 70/15/15 split of IQ chunks from continuous recordings (Sec. 7.1). Adjacent chunks share the same headset, user position, antenna placement, room RF background, and time period, so a random split leaks session context into the training set and can produce near-perfect accuracy even if the model has not learned app-specific features. The text does not state the number of participants or sessions, and Sec. 8.2 later restricts the attack to a single targeted user. Please add leave-one-session-out and leave-one-user-out evaluations, specify the collection timeline, and report per-session accuracy with confidence intervals.","section":"Sec. 6 / Sec. 7.1"},{"comment":"The case study claims a concealed-setup accuracy of about 0.99, but it uses the model 'well-trained in the subsection 7.1' on data whose temporal, room, and session relationship to the case-study recording is not reported. If the training set includes chunks from the same session or the same user position, the case study is not an independent validation of generalization. Please train the model only on data collected before the case-study session, or explicitly demonstrate that no training chunk overlaps the case-study recording.","section":"Sec. 7.2"},{"comment":"Fig. 30 shows USNR below 1 dB at several orientations (e.g., 0°, 180°, 225°, 315°), yet Fig. 31 reports app identification and activity recognition accuracy near 0.96 at those same orientations. If the emanation spikes are below the noise floor after averaging, the classifier must be relying on some stable cue other than the claimed app/activity emanation signature; the most parsimonious candidate is recording context shared between training and test chunks. This tension needs a control experiment, such as testing at a low-USNR orientation on data from a different session, and an explanation of which feature drives classification there.","section":"Sec. 7.3.6 / Figs. 30-31"},{"comment":"The cross-headset experiment trains and evaluates a separate model for each headset, so it does not test cross-hardware transfer. The statement in that section that the subtraction method can 'eliminate the hardware-dependent artifacts' is therefore not supported by the reported experiment. Please include a transfer test (train on one headset and test on the other for overlapping apps) or explicitly restate the claim as device-specific performance.","section":"Sec. 7.3.1"}],"minor_comments":[{"comment":"The averaging-gain calculation is incorrect: (14.7874-14.4809)/14.4809 is approximately 0.021, not 0.2, and expressing this as 'dB per second' is dimensionally inconsistent.","section":"Sec. 5.1 / Fig. 14"},{"comment":"The text says 'rewrite the spectrum expression' but Eq. (3) is a time-domain sum of sinusoids; please correct the wording and specify the summation bounds.","section":"Sec. 5.2 / Eq. (3)"},{"comment":"The app list contains 'Slupies' while Fig. 20 labels the same app 'Slurpies'; also Sec. 7.1 says 'beakroom' instead of 'break room'.","section":"Sec. 6 / Fig. 20 / Sec. 7.1"},{"comment":"The method paragraph contains an unresolved reference 'as shown in Fig. ??'; please fix the citation.","section":"Sec. 7.3.5"},{"comment":"The obfuscation countermeasure is evaluated only on simulated square waves (Figs. 32-33) with no description of the simulation or the daemon; please label it as a proof-of-concept and provide the simulation parameters.","section":"Sec. 8.1"},{"comment":"The four activities (entering, configuring, running, exiting) are used as labels throughout the evaluation but are never formally defined in terms of user actions or app states; adding precise definitions would support reproducibility.","section":"Sec. 6 / Sec. 7"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the manuscript's core idea is plausible and timely, and the main failure mode is experimental validation rather than theoretical inconsistency. The authors should be asked to supply cross-session/cross-user evaluation, confidence intervals, and a control for the low-USNR orientation result before the paper can be considered for acceptance. I do not see grounds for a novelty objection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is the first EM-side-channel attack on VR app identity and activity, and the core observation is plausible. The evaluation, however, does not support the 99% claim as stated.\n\nWhat's genuinely new: nobody has used electromagnetic emanations to distinguish VR apps and activities before. The paper identifies a real gap in VR side-channel work, and the small examples (Figs. 15-17) do show distinct spectra for different apps and activities. The engineering effort is real: substantial IQ data (7.5G and 42.5G) collected with a USRP N210, two commercial headsets, two rooms, 15 apps, and 4 activities, plus comparisons against LSTM and Transformer baselines. The signal-processing chain (noise-floor smoothing, interference subtraction, averaging FFT/STFT) is standard but appropriate for cleaning up weak emanations.\n\nThe soft spots are load-bearing. The biggest is the evaluation protocol. Chunks from a continuous recording are randomly split 70/15/15 into train/validation/test. Chunks recorded in the same session share the same headset, user position, antenna placement, room RF background, and time period, so the random split leaks session context. The 99% accuracy could largely reflect the model identifying the recording context rather than an app-specific fingerprint. No cross-session or cross-user test is reported, and Section 8.2 explicitly limits the work to a single targeted user. The stress-test note is right: the orientation result in Fig. 31 makes the problem concrete. Accuracy stays near 0.96 even at orientations where Fig. 30 shows USNR below 1 dB. If the emanation spikes are below the noise floor, the classifier must be using some other stable cue, and recording context is the obvious candidate.\n\nThere is also a quantitative error in the emanation enhancement claim. In Section 5.1, the USNR gain from averaging over 0.1s to 0.2s is computed as (14.7874-14.4809)/14.4809, which is about 2%, not 0.2 dB per second, and extrapolating that to a 2 dB gain over 10 seconds is unsupported. This is a minor-to-moderate flaw, but it undermines a specific enhancement claim.\n\nI would not desk-reject this. The research direction is timely, and the authors have shown a credible phenomenon. But the paper needs major revision before it is publishable: cross-session and cross-user validation, confidence intervals or repeated trials, and a corrected enhancement analysis. A serious referee should engage with it and ask for exactly those things; the topic deserves referee time, even though my own verdict on the current evidence is reject.","headline":"First EM side-channel on VR headsets with a credible core observation, but the 99% accuracy is built on same-session random splits that leak session context, so the generalization claim is unproven.","tokens_in":21799,"tokens_out":1741,"would_cite":false,"duration_ms":20687,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VReaves uses electromagnetic emanations from a VR headset to identify which of 15 apps a user is running and which of four activities they are performing, with 99% reported accuracy on both tasks.","keywords":["electromagnetic emanations","side channel","virtual reality","app identification","activity recognition","software-defined radio","spectrogram classification","ResNet"],"falsifier":"Collect emanation traces from the same VR apps on different days, in different rooms, and with different users, then train VReaves on one session and test it on another; if cross-session accuracy falls well below the reported 99%, the model was reading session-specific artifacts rather than a stable app-computation fingerprint.","tokens_in":20781,"feed_emoji":"📡","tokens_out":4345,"duration_ms":48862,"temperature":0.7,"pith_summary":"VReaves tries to establish that a passive radio receiver placed one to two meters from a VR user can learn what virtual-reality app the user is running and which of four app activities (entering, configuring, running, exiting) is happening, purely from electromagnetic emanations the headset emits unintentionally. The paper reports 99% accuracy for both tasks across commercial headsets, the Meta Quest 3 and HTC VIVE XR Elite. The stakes are immediate: if true, any nearby listener with a software-defined radio and a directional antenna can profile a VR user's behavior without malware, network access, or physical contact. The broader point is that the headset's embedded sensors turn it into a transmitter of computational activity that can be read at a distance.","feed_headline":"99% accurate: VR headset radio hum reveals app and activity","feed_subtitle":"Passively sniffed electromagnetic emissions identify 15 VR apps and four user actions, no malware or access needed.","key_machinery":"The load-bearing mechanism is a signal-processing pipeline that turns raw IQ samples into clean spectral images: movmedian noise-floor smoothing across scanned 10 MHz bands below 1 GHz, sliding-window spectrum subtraction to suppress quasi-stable ambient wireless signals, and non-coherent averaging of FFTs over time to lift weak emanation spikes above the noise. The classifiers are fine-tuned pre-trained ResNet18 networks, one fed with concatenated FFT outputs from multiple frequency bands for app identity and one fed with STFT spectrograms for activity recognition. The physical premise is that computational activity couples with clock signals through hardware components to produce emanation spikes, so the novel step is treating the entire headset as a multi-source radiator and letting a convolutional network separate the interleaved sources.","core_discovery":"The paper claims that the electromagnetic emanations from a VR headset are amplitude-modulated clock signals whose frequency spectrum and time-frequency spectrogram encode the computational activity of the camera, display, microphone, radio, and memory subsystems. Because different apps impose different computational patterns, the averaged FFT profiles of the emanations differ enough that a fine-tuned ResNet can distinguish fifteen VR apps; because activities change the temporal structure of the emanations, STFT spectrograms allow a second ResNet to distinguish entering, configuring, running, and exiting. The reported result is 99% accuracy for app identification and 99% accuracy for activity recognition, with accuracy staying near that level across tested distances, orientations, frequency bands, and two headset models.","pith_inferences":["This editorially inferred caveat is central: the evaluation uses random splits of a single data-collection campaign, so the 99% accuracy may partly reflect session-specific artifacts such as room RF background, user body position, or headset placement; a cross-session and cross-user test is the immediate test of generality.","The accuracy differences are likely tied to rendering and sensor workload, so apps with similar graphics load or similar sensor usage may be harder to separate than the fifteen-app set suggests.","The paper's proposed obfuscation countermeasure implies a testable defensive extension: running a daemon with randomized computational activity should measurably degrade the FFT and STFT separability of app classes.","If this result generalizes, app stores and VR platforms face a new privacy trade-off: immersive apps that use more sensors and rendering produce stronger, more identifiable emanation fingerprints."],"forward_implications":["A nearby attacker can silently track which VR app a person uses and when they enter, configure, run, or exit it, without installing anything on the victim's device.","The same pipeline could plausibly be applied to other head-mounted displays and wearable devices, since the emanation sources are generic hardware components rather than app-specific code paths.","Defenders cannot rely on shielding or jamming alone, because emanations are emitted automatically from multiple sensors, so countermeasures would need to alter the computational activity itself or obfuscate the emanation spectrum.","Activity recognition can expose behavioral structure, such as how long a user spends configuring versus playing, which feeds into inferences about the user's habits and personality.","If the relationship between emanations and computational activity holds, the attack may extend beyond app identity to video content and scene reconstruction, as the paper explicitly proposes as future work."],"supporting_citations":[{"why":"Supplies the emanation-source state detection technique used to distinguish active VR use from idle before app identification.","marker":"[53]"},{"why":"Provides the prior passive emanation characterization approach and the observation that ambient spectrum is quasi-stable, which the subtraction step relies on.","marker":"[52]"},{"why":"Establishes that camera and display emanations carry computational content, the basis for expecting VR app information in the spectrum.","marker":"[28]"},{"why":"Supplies the pre-trained ResNet architecture that VReaves fine-tunes for both FFT-based app identification and STFT-based activity recognition.","marker":"[54]"},{"why":"Defines the focused app set and the four app activities (entering, configuring, running, exiting) that the evaluation uses.","marker":"[37]"},{"why":"Defines the unintentional signal-to-noise ratio (USNR) metric used to measure emanation strength across antennas and distances.","marker":"[8]"}],"fun_headline_variants":["VR headset EM leaks reveal apps and actions at 99% accuracy","Electromagnetic side channel exposes VR app use and activity","Sniffing VR headset radio hum identifies apps and gestures","99% accurate: VR headset EM fingerprinting without malware","Passive EM sniffing decodes VR apps and user actions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The emanation pattern assigned to each app and activity is stable across users, sessions, and environments, yet the evaluation trains and tests on random splits of one data-collection campaign, so the reported 99% accuracy may reflect session-specific artifacts rather than the app's computational fingerprint alone.","fun_headline_variants_meta":{"raw":{"variants":["VR headset EM leaks reveal apps and actions at 99% accuracy","Electromagnetic side channel exposes VR app use and activity","Sniffing VR headset radio hum identifies apps and gestures","99% accurate: VR headset EM fingerprinting without malware","Passive EM sniffing decodes VR apps and user actions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000154,"raw_usage":{"total_tokens":1176,"prompt_tokens":879,"completion_tokens":297,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":211}},"tokens_in":495,"tokens_out":297,"duration_ms":3432,"temperature":1.0,"reasoning_tokens":211,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:31:09.780993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect emanation traces from the same VR apps on different days, in different rooms, and with different users, then train VReaves on one session and test it on another; if cross-session accuracy falls well below the reported 99%, the model was reading session-specific artifacts rather than a stable app-computation fingerprint.","supporting_citations":[{"cited_title":"On the Feasibility of Reasoning about the Internal States of Blackbox IoT Devices Using Side-Channel Information","cited_arxiv_id":"2311.13761","evidence_quote":"Supplies the emanation-source state detection technique used to distinguish active VR use from idle before app identification."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior passive emanation characterization approach and the observation that ambient spectrum is quasi-stable, which the subtraction step relies on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that camera and display emanations carry computational content, the basis for expecting VR app information in the spectrum."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the focused app set and the four app activities (entering, configuring, running, exiting) that the evaluation uses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the unintentional signal-to-noise ratio (USNR) metric used to measure emanation strength across antennas and distances."}],"review_version":1}