{"id":"c55ee8e0-a1dc-45e0-8ab9-7868c3ff0225","arxiv_id":"2411.10964","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"The paper sketches a device-oriented hierarchical ROI encryption scheme for AR sharing, but it is a preliminary position piece without new measurements.","lead":"This paper proposes encrypting different parts of shared AR video with different strengths, based on object sensitivity and on the display device (projector, smartphone, or AR glasses). It argues that bitstream-level region-of-interest encryption from the authors' earlier ROSS system beats pixel-level encryption, but offers no new experiments to back the claim.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract claims an adaptive system that dynamically adjusts encryption by display method, but Sec 2.2 states device selection is manual, so the central claim is not supported by the described implementation.","rationale":"I read the paper as a preliminary design/position statement rather than a completed system. In good faith, the authors explicitly flag limitations and call it preliminary, so missing user studies alone would not be disqualifying. However, the abstract's central promise is an adaptive system that dynamically adjusts encryption intensity based on display method, and the methodology explicitly states selection is currently manual. This contradiction cannot be repaired by future experiments; it is a mismatch between the claim and the described artifact. The reader's weakest assumption (privacy risk ordering) is also real and self-admittedly speculative, but it concerns the target of the hierarchy rather than whether the hierarchy exists. I therefore focus on the adaptive/manual inconsistency as the most load-bearing concern, while agreeing with the overall rejection.","tokens_in":3474,"tokens_out":5099,"duration_ms":56556,"concrete_test":"Conduct a behavioral test of the prototype: start an AR sharing session on a smartphone, then switch the output to a projector mid-session with no manual reconfiguration; observe whether the encryption level for the same ROI changes automatically. If it does not, the system lacks the adaptive behavior claimed in the abstract, and the contribution should be reframed as manual device-oriented encryption-level selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim: 'Our adaptive system dynamically adjusts the encryption intensity based on the AR display method.' Section 2.2 ('AR Display Device Identification and Switching') says: 'Currently, we manually determine the appropriate AR device for display and select the corresponding encryption level prior to each encryption.' Section 4 lists as future work: 'we aim to develop a fully automated end-to-end system by integrating additional components.' Therefore the system as described is not adaptive—it is a manual device-specific preset. If the display method changes during a session or the user does not manually preselect the level, the promised dynamic adjustment does not exist. This is an internal inconsistency, not merely a missing experiment: the central claim in the abstract and the implemented methodology describe different capabilities. The efficiency advantage (500MB vs 200KB, Sec 3) is also taken from the prior ROSS work [7] and not validated in the AR context, but the adaptive/manual mismatch alone breaks the paper's headline promise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a device-oriented hierarchical video encryption scheme for AR content sharing, in which the privacy safety of three display types (projector/SAR, smartphone/HHD, AR glasses/HMD) is used to determine the encryption strength applied to regions of interest such as faces, display content, and ID cards. The system builds on the authors' prior ROSS bitstream-level ROI encryption system and claims improved efficiency and adaptive behavior. The manuscript reports no new experiments: Section 3 imports the key efficiency numbers from prior work [7], Section 2.2 states that device selection is performed manually, and Section 4 defers automation, user studies, and performance tests to future work.","tokens_in":3674,"tokens_out":4251,"duration_ms":47362,"significance":"Privacy protection in multi-device AR content sharing is a timely and practically important problem. Combining bitstream-level ROI encryption with device-dependent privacy levels is a reasonable direction, and the paper clearly identifies the trade-off between encryption safety and real-time performance. If the claims were validated, the approach could reduce encrypted data volume compared with pixel-level encryption while adapting protection to display context. However, none of the novel components are validated in this manuscript: the device-to-safety mapping is admittedly speculative, the semantic importance ordering is asserted without rationale, and the quantitative efficiency evidence is lifted from the authors' own prior work. The paper's main strength is its honest articulation of limitations, not empirical support for its central claims.","major_comments":[{"comment":"The abstract's central claim of an 'adaptive system [that] dynamically adjusts the encryption intensity based on the AR display method' is contradicted by the described implementation. Section 2.2 states that the device is 'manually determine[d]' and the encryption level is selected 'prior to each encryption,' and Section 4 lists a 'fully automated end-to-end system' as future work. The system as presented is therefore a manual per-device preset, not an adaptive or dynamic mechanism, so the paper's headline claim is not supported by its own methodology.","section":"Sec. 2.2 and Sec. 4"},{"comment":"The only quantitative result supporting the efficiency claim, the comparison between pixel-level encryption (500 MB) and bitstream-level ROSS encryption (200 KB) at PSNR ≈ 15 dB, is taken from the authors' prior paper [7] and is not reproduced or extended to the AR content-sharing context. Since this number is the sole evidence for the claim of a 'more efficient solution,' the central efficiency advantage rests entirely on a self-citation rather than on any measurement, trace, or benchmark presented here.","section":"Sec. 3"},{"comment":"The core input to the proposed system, the privacy safety ordering of display devices (projector lowest, smartphone balanced, AR glasses highest), is acknowledged in the text as 'a mere example and not a thorough risk assessment.' The paper also asserts without support or justification that ID cards are 'slightly important' while display content is 'moderately important,' which is counterintuitive given the sensitivity of ID cards in AR backgrounds. Because the hierarchical encryption levels are determined by these orderings, an incorrect ordering would mis-target the actual privacy threats; no threat model, user study, or domain analysis is provided to substantiate either ordering.","section":"Sec. 3"},{"comment":"No experiments, user study, or performance test are reported, and Section 4 explicitly states that user studies and performance tests are future work. Consequently, the stated goal of balancing 'encryption safety' and 'real-time performance' is never evaluated, and there is no evidence that the proposed system is usable in a realistic AR pipeline. The manuscript is a position statement rather than a validated system paper, which leaves the central claims unsupported.","section":"Secs. 1 and 4"}],"minor_comments":[{"comment":"The phrase 'title number: 12×16' appears to be a typo for 'tile number: 12×16'; please correct it and define the tile configuration in the text.","section":"Fig. 2 caption"},{"comment":"The acronyms SAR, HHD, and HMD are used in the figure caption but are not defined in the text where they first appear; please define them at first use.","section":"Fig. 1 and Sec. 2"},{"comment":"Reference [7] is cited as the source for the ROSS system, but the reference title is 'SDM: Semantic Distortion Measurement for Video Encryption'; the relationship between ROSS and SDM should be stated explicitly so readers can locate the system description.","section":"Sec. 2.1"},{"comment":"The mapping from device privacy safety level to encryption level is described only qualitatively; an explicit table or formula showing how the safety level determines concrete encryption parameters (e.g., QP, tile configuration, or encryption strength) would make the proposal easier to test and reproduce.","section":"Sec. 3"}],"recommendation":"reject","confidential_remarks":"This manuscript appears to be an extended abstract or position paper rather than a full research contribution. The quantitative efficiency claim is a self-citation from the authors' prior work [7], the adaptive behavior promised in the abstract is contradicted by the manual selection described in Section 2.2, and the privacy ordering is explicitly unvalidated. Revisions would require substantial new experiments and a reframing of the contribution; as it stands, the central claims are not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the kernel idea—basing ROI encryption strength on which AR display receives the shared video—is a reasonable, under-explored direction. The authors are also upfront that their device-privacy ordering is speculation, not a risk assessment. But the paper does not deliver what the abstract promises. The abstract says the system is adaptive and dynamically adjusts encryption intensity, while Section 2.2 says device selection and encryption level are chosen manually before each encryption. That is not a missing experiment; it is a mismatch between the central claim and the described implementation. The Limitations section lists automation as future work, which compounds the problem.\n\nWhat is genuinely new: the three-tier mapping from display type (projector, smartphone, AR glasses) to a privacy safety level, and the suggestion to vary semantic-object encryption accordingly. Nothing in the cited literature does exactly that. The paper also positions itself competently within AR privacy and ROI encryption work, and it explicitly flags the risk ordering as a guess. Those are good-faith moves.\n\nThe soft spots are large. The 500MB-to-200KB efficiency figure comes entirely from the authors' prior ROSS paper, not from any experiment here, so it adds no new evidence. The privacy-safety ordering is asserted without user study or contextual analysis, as the authors admit. And the adaptive claim is contradicted by the manual selection description. Read generously, this is a well-scoped suggestion for future work; read literally, it is an overclaim.\n\nWho is this for? A workshop audience working on AR privacy would get a conversation starter. A reader looking for a validated technique will be disappointed. I would not cite it in its current form, and I would not send it to a serious referee until the abstract is aligned with the implemented system and at least one small demonstration is added—even a simulated encryption-level switch across the three device types would help. As is, it is a sketch, not a contribution.","headline":"A plausible design sketch that overstays its abstract: the device-to-privacy mapping is new but manual, and the efficiency numbers are inherited, so the current paper does not substantiate its headline claim.","tokens_in":4162,"tokens_out":2674,"would_cite":false,"duration_ms":28353,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that bitstream-level ROI video encryption can protect the physical background in AR content sharing far more efficiently than pixel-level encryption, with encryption intensity tuned to the exposing display device.","keywords":["augmented reality","visual privacy","region of interest encryption","bitstream-level encryption","hierarchical privacy","HEVC","selective encryption","AR content sharing"],"falsifier":"A controlled measurement showing that AR-glasses wearers' content is captured by nearby bystanders at rates equal to or higher than projector viewers would falsify the device-safety ordering, as would a benchmark where bitstream-level ROI encryption on real AR backgrounds yields significantly worse than the claimed 15 dB PSNR at 200 KB.","tokens_in":3298,"feed_emoji":"🔐","tokens_out":5976,"duration_ms":57425,"temperature":0.7,"pith_summary":"This paper proposes that private details in the physical background of AR content sharing—faces, ID cards, display content—can be protected by encrypting only the regions of interest inside the compressed video stream, rather than the whole frame at the pixel level. It argues this bitstream-level approach is dramatically more efficient, citing a comparison where the same visual distortion (PSNR around 15 dB) costs 500 MB with pixel-level encryption versus 200 KB with the ROSS bitstream system. The encryption strength is then adjusted according to which AR display is receiving the content: projection is treated as the most exposed, smartphones as intermediate, and AR glasses as the most private. If correct, this gives AR collaboration a way to balance privacy protection against the real-time performance demands of augmented reality.","feed_headline":"Bitstream AR encryption cuts privacy overhead from 500MB to 200KB","feed_subtitle":"Encryption strength adapts to the display: strongest for projectors, lightest for AR glasses, keeping sharing real-time.","key_machinery":"The load-bearing mechanism is the Region Of Semantic Saliency (ROSS) encryption system, an ROI video encryption variant that algorithmically encrypts foreground objects at the bitstream level instead of requiring user selection. It works through the tile mechanism of the HEVC encoder, which codes distinct video areas independently so that only tiles covering sensitive objects are encrypted. The paper couples this with a semantic-importance classification—faces are highly important, displayed content moderately important, ID cards slightly important—and a device-safety ordering, so the encryption level is chosen per display device. This combination is what lets the system claim to conserve code stream and reduce encryption overhead while still protecting the physical background.","core_discovery":"The paper's central claim is that ROI video encryption, performed at the bitstream level during HEVC encoding, can be brought into AR content sharing to hide sensitive parts of the physical environment, and that the encryption intensity should scale with the privacy exposure of the target display. The authors treat projection as a public display with the lowest privacy safety level, smartphones as offering balanced control, and AR glasses as the most private because content is seen by a single wearer; more exposed devices receive stronger encryption. They report that at comparable compression quality (PSNR $\\approx 15$ dB), pixel-level encryption consumes 500 MB while the bitstream-level ROSS system requires only 200 KB, and they position this as making hierarchical, device-aware privacy protection feasible without breaking real-time AR performance. The paper is explicit that the device risk ordering is a speculative example rather than a thorough risk assessment, and that the current implementation still selects the AR display manually.","pith_inferences":["The 500 MB versus 200 KB figure comes from the underlying ROSS system on a standard test sequence, not from the AR hierarchy itself; an AR-specific benchmark would be needed to confirm the savings survive real background scenes and multiple tile partitions.","The assumed device ordering could be tested directly: measuring actual shoulder-surfing or bystander capture rates for projection, smartphone, and AR glasses in controlled settings would tell whether projector really is the riskiest.","A natural extension is to automate device detection and encryption-level switching using existing sensor or hardware identification methods, turning the manual selection into the end-to-end system the authors plan.","The semantic-importance categories are limited to three example objects; a larger taxonomy, such as license plates, medical records, or screens, would be needed before deployment."],"forward_implications":["Physical backgrounds in AR sharing can be protected without encrypting the whole frame, so common sensitive items such as faces and cards can be hidden at a tiny fraction of the data cost of pixel-level encryption.","Privacy protection can be made display-aware: a projector feed can be encrypted more heavily than a smartphone or AR-glasses feed, so each viewer gets the minimum protection their exposure requires.","Because encryption happens inside the HEVC encoding stage, the approach can preserve real-time AR performance better than post-hoc pixel encryption.","The semantic-importance ranking provides a natural way to allocate encryption effort: the most sensitive object types receive the strongest protection first."],"supporting_citations":[{"why":"Supplies the ROSS bitstream-level ROI encryption system and the semantic distortion measurement used as the core mechanism.","marker":"[7]"},{"why":"Establishes ROI encryption for HEVC coded video, providing the tile-based integration the paper builds on.","marker":"[4]"},{"why":"Shows a prior adaptation of encryption to XR displays, namely tracking-tolerant visual cryptography, which the paper distinguishes from its AR-focused approach.","marker":"[3]"},{"why":"Documents the tradeoff between selective encryption and computational cost, motivating the balance between security and real-time AR performance.","marker":"[9]"},{"why":"Provides evidence of privacy risks in mobile AR applications, grounding the need for device-aware privacy protection.","marker":"[8]"}],"fun_headline_variants":["Bitstream ROI encryption shrinks AR privacy cost from 500MB to 200KB","Adaptive bitstream encryption tailors AR privacy to display exposure","Bitstream-level ROI encryption makes AR sharing privacy feasible in real time","Hierarchical bitstream encryption protects AR privacy with 2500x less data","Bitstream ROI encryption scales with display risk, from projector to glasses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire encryption-strength schedule rests on the unmeasured assumption that projectors are the most exposed displays, smartphones are intermediate, and AR glasses are the safest.","fun_headline_variants_meta":{"raw":{"variants":["Bitstream ROI encryption shrinks AR privacy cost from 500MB to 200KB","Adaptive bitstream encryption tailors AR privacy to display exposure","Bitstream-level ROI encryption makes AR sharing privacy feasible in real time","Hierarchical bitstream encryption protects AR privacy with 2500x less data","Bitstream ROI encryption scales with display risk, from projector to glasses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000479,"raw_usage":{"total_tokens":2345,"prompt_tokens":895,"completion_tokens":1450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":511,"completion_tokens_details":{"reasoning_tokens":1354}},"tokens_in":511,"tokens_out":1450,"duration_ms":12155,"temperature":1.0,"reasoning_tokens":1354,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:04:38.447977+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled measurement showing that AR-glasses wearers' content is captured by nearby bystanders at rates equal to or higher than projector viewers would falsify the device-safety ordering, as would a benchmark where bitstream-level ROI encryption on real AR backgrounds yields significantly worse than the claimed 15 dB PSNR at 200 KB.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ROSS bitstream-level ROI encryption system and the semantic distortion measurement used as the core mechanism."},{"cited_title":"Farajallah, W","cited_arxiv_id":null,"evidence_quote":"Establishes ROI encryption for HEVC coded video, providing the tile-based integration the paper builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows a prior adaptation of encryption to XR displays, namely tracking-tolerant visual cryptography, which the paper distinguishes from its AR-focused approach."},{"cited_title":"Massoudi, F","cited_arxiv_id":null,"evidence_quote":"Documents the tradeoff between selective encryption and computational cost, motivating the balance between security and real-time AR performance."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides evidence of privacy risks in mobile AR applications, grounding the need for device-aware privacy protection."}],"review_version":1}