{"id":"202021a9-935d-40c0-9bda-ce377a9803cd","arxiv_id":"2607.23984","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"44.62% of audited mobile-AR privacy policies omit more than eight of twenty-two U.S. state-law disclosure requirements, with rights and biometric-data disclosures among the most commonly missing.","lead":"This paper audits 4,116 English-language privacy policies of mobile augmented-reality apps against a taxonomy built from 20 U.S. state privacy laws, and reports that 44.62% miss more than eight disclosure requirements. It is the first large-scale privacy-policy audit of the MAR ecosystem under U.S. state laws, and ships a dataset, taxonomy, and automated pipeline for future compliance auditing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Over-inclusive legal taxonomy may inflate severe-disclosure rate: single-state requirements applied to all policies.","rationale":"The reader's weakest assumption identified both the over-inclusive taxonomy and the LLM under-labeling as threats. I focus on the legal-applicability issue because it is more fundamental: even a perfectly accurate annotator would over-count violations if the requirements themselves do not apply. The paper's own taxonomy (Section IV-B) explicitly notes that several requirements are unique to one or a few states, yet the audit applies them uniformly to all 4,116 policies. This directly inflates the number of missing requirements per policy and, by extension, the 44.62% severe-omission rate and the four >90% violation rates. The paper does not self-flag this limitation in Section VII-B. The concrete test—recomputing with a restricted taxonomy and/or statutory thresholds—would settle whether the headline numbers are robust. If the numbers drop substantially, the central claim is not sustainable as stated; if they do not, the concern is mitigated. The paper's pipeline validation and release of artifacts are valuable, but they do not address this external-validity issue. Thus the reader's CONDITIONAL verdict remains appropriate, and I would not change it without the additional analysis.","tokens_in":19743,"tokens_out":5957,"duration_ms":56016,"concrete_test":"Recompute the severe-omission percentage (44.62%) and per-requirement violation rates using a restricted taxonomy that includes only requirements common to a supermajority of states (e.g., required by at least 19 of 20) or that applies each state's statutory thresholds (e.g., exclude apps below California's $25M revenue / 100k consumers threshold from CCPA-specific requirements). If the restricted severe-omission rate drops materially below 44.62% or the four >90% violation rates fall below 90%, the original figures are an artifact of legal over-inclusion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central statistic—44.62% of audited policies with more than eight missing requirements—depends on counting violations of requirements that are not legally applicable to every app. The taxonomy in Section IV-B includes provisions found in only a few states: D2 (sources of PI) is required only by California, D6 by California and Minnesota, PA1 by four states, PA2 only by Oregon, PA3 by New Hampshire and Minnesota, R8/R9 by Minnesota and Maryland, and PC3/PC4 only by Colorado. The audit then treats every U.S.-available policy as obligated to satisfy all 22 requirements regardless of whether each state's law actually applies to the app. Many state comprehensive privacy laws contain applicability thresholds (e.g., revenue, number of consumers) that would exclude small or low-volume apps, and an app may not be targeted at every state. A policy that omits D2, D6, PA1, PA2, PA3, R8, R9, PC3, and PC4 could exceed the '>8 missing' threshold without violating any law that actually binds it. The paper performs no sensitivity analysis restricting the taxonomy to broadly applicable requirements, nor does it model statutory thresholds. Consequently, the headline violation rates may be substantially overstated.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a large-scale audit of mobile AR (MAR) app privacy policies against U.S. state comprehensive privacy laws. The authors build a Google Play MAR dataset (8,013 metadata records, 6,620 APKs, 6,426 privacy policies), derive a 22-requirement disclosure taxonomy from 20 state statutes, and apply a four-stage LLM pipeline (processing, annotation, extraction, normalization) to assess 4,116 English-language policies. They report that 44.62% of audited policies miss more than eight requirements, with particularly high violation rates for rights-related and biometric-specific disclosures. The dataset, taxonomy, and pipeline are released for reproducibility.","tokens_in":20073,"tokens_out":7077,"duration_ms":73601,"significance":"This is the first systematic attempt to audit MAR privacy policies against the fragmented U.S. state-law landscape. The contributions are substantial: a public MAR dataset, a law-derived taxonomy, and a largely traceable pipeline with human-validated labels and an extraction-fidelity check (98.32% at threshold 0.95). The descriptive ecosystem analysis (RQ1) and the taxonomy derivation (RQ2) are credible and valuable. However, the quantitative headline (RQ3) is not yet reliable because the taxonomy is a union of state-specific obligations applied uniformly to all policies, and because the pipeline's documented under-labeling would systematically inflate measured violations. With a sensitivity analysis restricted to broadly applicable requirements and error-bounded estimates, the paper could make a strong contribution.","major_comments":[{"comment":"The audit scores every U.S.-available policy against all 22 requirements even though the taxonomy is a union of obligations from 20 state laws, several of which appear in only one or a few states: D2 is required only by California, D6 by California and Minnesota, PA1 by four states, PA2 only by Oregon, PA3 by New Hampshire and Minnesota, R7 only by California, R8/R9 by Minnesota and Maryland, and PC3/PC4 only by Colorado. Many of these state laws also have applicability thresholds (revenue, consumer counts, processing volume) and require a state nexus, so a policy that omits D2, PA2, PC3, PC4, etc., may still comply with every law that actually binds it. The paper performs no sensitivity analysis restricting the taxonomy to requirements shared across states or modeling statutory thresholds. Consequently, the headline 44.62% severe-omission rate and the >90% violation rates (e.g., R7 at 9","section":"IV-B, VI-B"},{"comment":"The pipeline evaluation shows document-level multi-label Micro-F1 of 87.69% and Jaccard of 73.56%. The paper further reports that under-labeling dominates passage-level mismatches (1,183 of 1,999; 59.17%), with false negatives arising when passages refer only abstractly to 'information', 'data', or 'third parties'. Since a requirement is counted as violated when the pipeline finds no matching passage, systematic under-labeling directly inflates violation counts. The paper does not report document-level precision/recall and does not provide an error-bounded estimate of the effect of label error on the severe-omission rate. A best-case/worst-case analysis based on the confusion matrix, or a re-annotation of a random sample with corrected rates, is needed before the quantitative findings can be accepted.","section":"VII, Table VIII"},{"comment":"The conditional-chain violation rates are computed only over policies whose trigger is detected by the same pipeline (e.g., Chain 3 uses policies where D1 discloses biometric data). If trigger detection is conservative, the denominator is reduced and the conditional rates describe a subset that may not be representative of all apps engaging in the underlying practice. The paper should report how many policies were excluded due to absent triggers and, if feasible, validate the triggers against the human-annotated ground truth.","section":"VI-B3"}],"minor_comments":[{"comment":"Headline percentages such as 44.62% and the >90% violation rates are reported without confidence intervals; add them, especially given the n=3,855 denominator.","section":"Abstract, VI-B"},{"comment":"The legend symbols (blank, G#, ⊙, #) are visually hard to distinguish, especially after typesetting. Use explicit text labels such as 'Baseline', 'Triggered', 'Conditional', 'Not required'.","section":"Table IV"},{"comment":"Figure 3a labels C1–C4 as if they were individual requirements; the text calls them 'logic chains'. Clarify the terminology so the violation percentages are not misinterpreted as statutory requirements.","section":"Fig. 3a, VI-B"},{"comment":"The text moves from 6,426 collected policies to 4,116 English policies but does not state how many were excluded by language detection or by the 261 'unlabeled' files. Report these numbers explicitly and discuss potential language bias.","section":"III-F, V-A"},{"comment":"The relationship between '1,529 annotations' in Section VI-C2 and '1,999 mismatches' in Section VII is confusing. Define which set each number refers to.","section":"VI-C, VII"},{"comment":"The secondary APPG-analysis (766/6,426 policies) is introduced without methodological detail on the keyword fingerprinting. Either add a brief description or move the details to the supplemental materials.","section":"VII-A"}],"recommendation":"major_revision","confidential_remarks":"The paper has a solid, reproducible pipeline and a valuable new dataset, but the central quantitative claims are not yet defensible. The two load-bearing issues are (i) the over-inclusive union taxonomy applied to all policies and (ii) the systematically biasing under-labeling, both of which inflate the measured disclosure gaps. These are fixable within the manuscript's scope via sensitivity analyses and error-bounded estimates, so I recommend major revision rather than rejection. I also note that the full dataset is only released upon acceptance, which limits independent verification during review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read. This is the first systematic audit of MAR privacy policies against the current U.S. state-law patchwork, and it ships a genuinely reusable dataset, taxonomy, and pipeline. But the headline statistic—44.62% of policies missing more than eight requirements—is likely inflated, because the audit treats all 22 taxonomy requirements as binding on every U.S.-available app even though several (D2, D6, PA1, PA2, PA3, R8, R9, PC3, PC4) appear in only one to four states. The paper never models statutory applicability thresholds or app-level state targeting, and it runs no sensitivity analysis restricted to broadly applicable requirements. A policy could exceed the '>8 missing' threshold while violating no law that actually binds it. This is a real flaw, not a quibble.\n\nThe positive core is solid. The MAR dataset construction is careful—multi-stage verification with human review, BFS expansion, and a fair evaluation of the classifier. The 22-requirement taxonomy is a useful consolidation of 20 state laws, and the conditional logic chains are a genuine advance over flat checklists. The pipeline produces traceable quotes, and the authors validate extraction fidelity (98%+ at 0.95 similarity) and audit judgments against a gold standard with substantial inter-rater agreement. They also disclose the dominant error mode: under-labeling (59% of mismatches), with document-level recall around 0.88. That means the measured violations are probably overestimated, consistent with the legal-applicability concern. No confidence intervals on the headline rates either, so treat the exact numbers as upper bounds.\n\nWho gets value: anyone working on privacy-policy compliance measurement, app-store auditing, or LLM-based policy analysis. The dataset and taxonomy are contributions in themselves, and the limitations section is honest about static-policy-only scope.\n\nRecommendation: send it to peer review, but ask for revision. The authors should rerun with a restricted taxonomy (e.g., requirements shared by at least N states), model applicability thresholds, and report precision/recall-adjusted bounds. If the headline survives that, it's a strong empirical result; if not, the paper remains useful but with a more measured conclusion.","headline":"First large-scale MAR audit under U.S. state privacy laws, with a strong dataset/pipeline, but the 44.62% severe-disclosure headline is likely inflated by an over-inclusive legal taxonomy.","tokens_in":20479,"tokens_out":3620,"would_cite":true,"duration_ms":34400,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 4,116-policy audit finds mobile AR apps miss most U.S. state-law privacy disclosures.","keywords":["mobile augmented reality","privacy policy audit","U.S. state privacy laws","disclosure taxonomy","LLM-based auditing","biometric data","Google Play","compliance measurement"],"falsifier":"Have a lawyer, working without the pipeline, audit a random sample of 100 of the 4,116 policies against the statutes of the specific states where each app is distributed, applying the same 22-requirement taxonomy by hand. If the human audit finds violation rates substantially below the paper's figures—particularly for the four requirements with violation rates above 90%—the automated numbers would not reflect real legal exposure.","tokens_in":19716,"feed_emoji":"📱","tokens_out":4126,"duration_ms":35383,"temperature":0.7,"pith_summary":"The paper sets out to measure whether mobile augmented reality (MAR) apps—which collect camera, location, spatial, and biometric-adjacent data—disclose their data practices well enough to satisfy the privacy laws now in effect in 20 U.S. states. It builds a dataset of 8,013 MAR apps from the Google Play store, distills the state laws into a 22-requirement disclosure taxonomy, and runs an LLM-based pipeline over 4,116 English-language privacy policies. The central finding is that disclosures are systematically incomplete: 44.62% of audited policies miss more than eight requirements, and four requirements—including the right to non-discrimination and an appeal mechanism—are missing in over 90% of policies. The gaps are largest exactly where MAR risk is highest: biometric-data collection is disclosed in 601 policies, but nearly all of those policies lack the incident-response and deletion disclosures that Colorado law requires. A sympathetic reader would take this as evidence that MAR privacy policies have not kept pace with the fragmented, evolving U.S. state-law landscape.","feed_headline":"44% of mobile AR apps omit eight privacy disclosures","feed_subtitle":"Audit of 4,116 Android AR privacy policies against 20 U.S. state laws finds rights and biometric gaps.","key_machinery":"The carrying object is the auditable disclosure taxonomy: a consolidated set of 22 requirements (Data Transparency D1–D6, Rights Notice R1–R9, Privacy Controls PC1–PC4, Policy Administration PA1–PA3) derived from the 20 state laws, with applicability modeled as baseline, triggered, or conditional. Four conditional logic chains encode legal preconditions—e.g., if a policy discloses biometric-data collection, it must also disclose an incident-response protocol and deletion guidelines; if it discloses selling or sharing, it must provide an opt-out. The second key mechanism is a four-stage automated pipeline (clean, annotate, extract, normalize) that labels passages, quotes exact evidence spans,","core_discovery":"The paper's central claim is that MAR privacy policies, at ecosystem scale, fail to reflect the disclosure obligations imposed by U.S. state comprehensive privacy laws. Using a taxonomy of 22 requirements derived from 20 state statutes—organized into 5 baseline requirements, 10 triggered requirements, and 7 conditional requirements in 4 logic chains—the audit finds that only a minority of policies are complete. More than half omit core consumer rights such as the right to access (58.1%), correction (59.5%), and deletion (50.8%); 97.4% lack an appeal mechanism; and among policies that disclose biometric-data collection, 99.8% lack a biometric incident-response protocol and 97.2% lack biometri","pith_inferences":["The audit measures what policies say, not what apps do; if a policy is incomplete, actual runtime behavior might be even less protective, so the findings are likely a lower bound on the transparency problem but say nothing about data-practice compliance.","The dominant under-labeling error mode (59% of mismatches are false negatives) means the measured violation rates could be overestimated; a more lenient re-labeling of ambiguous 'data' or 'third parties' passages would likely reduce—but probably not eliminate—the headline gaps.","Because the taxonomy unions all 20 state laws, a policy is counted as violating a requirement even if it targets a state that does not impose it (e.g., a biometric incident protocol required only by Colorado); a state-specific audit would yield lower but still substantial violation rates.","The same pipeline could be turned around to test whether GDPR-oriented policies, which are structurally similar in many MAR apps, are any closer to satisfying U.S. state-law expectations; the paper hints at regime mismatch when policies are written in GDPR terminology."],"forward_implications":["If the findings hold, regulators have concrete evidence that MAR-specific sensitive data, such as spatial maps and bystander imagery, falls outside current definitions of sensitive personal information and needs explicit statutory treatment.","App markets could deploy automated screening tools to flag policies missing required disclosures during pre-release or periodic review, rather than relying on developer self-certification.","Developers cannot rely on generic privacy-policy templates: the paper's secondary analysis finds template-generated policies still show 5–8 violations in 56% of cases and more than 8 in 24% of cases.","The released dataset, taxonomy, and pipeline let other researchers reproduce the audit and extend it to other app categories or to future state laws.","The policy-maintenance gap implies that even apps with reasonable policies fall out of compliance when apps are updated but policies are not; 23.90% of U.S. MAR apps had not been updated since 2023 or earlier."],"fun_headline_variants":["AR apps fail state privacy laws: 45% omit 8+ disclosures","97% of AR apps lack a privacy appeal process","99.8% of AR apps lack biometric incident response","State privacy audit: AR policies show systemic gaps","Mobile AR privacy policies understate state law duties"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The audit assumes the union of 22 disclosure requirements drawn from all 20 state laws applies to every U.S.-available app, and that a requirement is failed whenever the automated pipeline finds no matching passage; if the legal scope is too broad or the pipeline under-labels, the measured gaps are inflated.","fun_headline_variants_meta":{"raw":{"variants":["AR apps fail state privacy laws: 45% omit 8+ disclosures","97% of AR apps lack a privacy appeal process","99.8% of AR apps lack biometric incident response","State privacy audit: AR policies show systemic gaps","Mobile AR privacy policies understate state law duties"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2447,"prompt_tokens":816,"completion_tokens":1631,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":560,"completion_tokens_details":{"reasoning_tokens":1551}},"tokens_in":560,"tokens_out":1631,"duration_ms":13511,"temperature":1.0,"reasoning_tokens":1551,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:19:30.458960+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a lawyer, working without the pipeline, audit a random sample of 100 of the 4,116 policies against the statutes of the specific states where each app is distributed, applying the same 22-requirement taxonomy by hand. If the human audit finds violation rates substantially below the paper's figures—particularly for the four requirements with violation rates above 90%—the automated numbers would not reflect real legal exposure.","supporting_citations":[],"review_version":1}