{"id":"70b58f39-1dc2-4f27-884f-26feed447c05","arxiv_id":"2501.19173","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Many prior studies that apply Contextual Integrity to language models omit the theory's core tenets, which can make their privacy conclusions unreliable.","lead":"This position paper argues that many researchers who use the privacy theory Contextual Integrity to study large language models are applying it too loosely, leaving out core parts of the theory. The authors lay out the theory's four key requirements, score nine prior studies against them, and warn that prompt sensitivity and answer-order bias can make privacy evaluations unreliable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"T2's categorical exclusion of legal and crowdsourced norm sources is the load-bearing step; if those can be evidence of norms, the all-✗ Table 1 pattern collapses and the central claim overgeneralizes.","rationale":"The strongest claim is a universal negative about the LLM-CI literature: existing works do not embrace CI's fundamental tenets, so their privacy conclusions are unreliable. The load-bearing step is therefore not the survey itself but the normative standard used to code 'deviation.' The paper asserts that all four tenets, as the authors enumerate them, must be satisfied, and in §3 it makes two categorical exclusions: legal statutes and crowdsourced preferences cannot be privacy norms. These exclusions do the work of producing the all-✗ T2 and T4 columns in Table 1. But CI scholarship routinely uses law and empirical surveys as sources or indicators of norms; Nissenbaum's own co-authored Martin & Nissenbaum (2016) uses survey vignettes, and the paper's Table 2 concedes governing institutions can reflect norms. The paper never provides a criterion for distinguishing a 'proxy' from 'evidence,' so it cannot justify marking a work that grounds HIPAA-based flows in a health context as categorically non-CI. Without that categorical step, the summarized verdict 'all prior works ✗ for T2' loses its basis, and the central claim reduces to 'some works borrow CI vocabulary without engaging its normative apparatus'—a useful but narrower position. This is why I keep the reader's CONDITIONAL verdict: the critique is salvageable with a more careful standard, but as written the central claim overgeneralizes. The concrete re-coding test would show whether the all-✗ pattern is an artifact of the strict criterion. I do not see a more serious internal inconsistency; the experimental-hygiene section is ancillary and the sample-selection concern is secondary to the definitional one.","tokens_in":14843,"tokens_out":4603,"duration_ms":47038,"concrete_test":"Run a re-coding sensitivity analysis on Table 1. For each of the nine papers, extract the sentence(s) where they define or source 'privacy norms' (e.g., the HIPAA-derived ontology in Fan et al./Li et al.; the vignette survey in Mireshghallah et al./Shao et al.). Re-code T2 as aligned whenever the paper uses a legal or crowdsourced source as evidence of norms for a stated context, rather than as an unexamined ground-truth definition. If this re-coding changes any of the nine T2 marks (or any T4 marks), Table 1's all-✗ pattern is not robust; the 'inadequate application' claim should be softened to 'some applications use proxies without justification.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central charge depends on treating its four-tenet checklist as the only legitimate way to use CI, and specifically on the categorical claims in §3 that 'Legal Statutes are not Privacy Norms' and 'Crowdsourced Preferences are not Privacy Norms.' Both claims are stronger than the paper needs and sit in tension with its own Table 2 position that governing institutions can be designed to support and reflect privacy norms. Law can encode settled contextual norms (e.g., HIPAA in health contexts), and structured surveys are the standard empirical method for eliciting norms in CI research (Martin & Nissenbaum 2016 is itself a survey instrument). The paper distinguishes 'evidence of norms' from 'ground truth' only rhetorically; it then marks every audited work ✗ on T2 because they use legal statutes or preferences. If those sources can legitimately function as evidence of norms, the blanket all-✗ T2 row is an overreach, and the inference from 'these works use proxies' to 'these works inadequately apply CI' is unsupported. The same problem propagates to T4, where deferring to laws/preferences is treated as automatically failing the CI heuristic rather than as a possible input to it. Because the conclusion rests on this binary, the central claim is conditional on a contestable standard.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position paper argues that recent LLM privacy evaluations that invoke Contextual Integrity (CI) do not actually apply the theory, because they fail to uphold its four tenets (T1-T4). The authors clarify the tenets, survey nine prior works, tabulate deviations in Table 1, and add a call for experimental hygiene around prompt sensitivity and position bias. Their central claim is that inadequate application risks incorrect conclusions and flawed privacy-preserving designs.","tokens_in":15086,"tokens_out":4831,"duration_ms":46389,"significance":"The paper makes a timely and useful corrective: it gives the ML community a compact statement of CI's core tenets, and Table 1 is a clear systematization that will help readers see where partial uses diverge. Section 4's emphasis on prompt sensitivity and position bias in CI-based vignette studies is a concrete, actionable contribution. The force of the central claim, however, is conditional on a strict standard for what counts as legitimate norm identification and normative assessment; that standard is currently stated too categorically in Section 3.","major_comments":[{"comment":"The categorical claim that legal statutes are not privacy norms is overbroad and is internally inconsistent with the paper's own position in Table 2, which states that 'Governing institutions can be designed to support and reflect privacy norms.' This concession implies that legal instruments can encode settled contextual expectations, as HIPAA does for health contexts. The text, however, uses the categorical claim to mark all surveyed works as ✗ on T2 merely because they draw on legal statutes or crowdsourced preferences. This conflates 'source or evidence of a norm' with 'ground truth.' If these sources can legitimately serve as evidence of norms, the all-✗ T2 row overreaches, and the inference from 'uses proxies' to 'inadequately applies CI' is unsupported. The authors should restate T2 as requiring that norm sources be validated and subjected to normative assessment, rather than prohibiting legal or survey-based inputs outright.","section":"Section 3, 'Legal Statutes are not Privacy Norms' and Table 2"},{"comment":"The blanket ✗ for T4 is based on the observation that the surveyed works 'defer to laws, policies, and collective preferences as proxies' for ethical legitimacy. This is a non sequitur: using legal or survey input as evidence within a CI heuristic, for example at Level 1 where interests and preferences are examined, is not the same as treating that input as final ethical authority. The paper does not show that any of the nine works considered and rejected the heuristic; it shows that they did not mention it. That absence is a valid observation, but the inference to 'failure' requires a standard that distinguishes omission from affirmative wrong treatment. Without that distinction, the T4 column overstates the case.","section":"Section 3, T4: CI Heuristic"},{"comment":"Mireshghallah et al. (2023) are marked ✗ for using only three CI parameters, yet the paper itself acknowledges that Martin & Nissenbaum (2016), the authoritative source for the survey template they follow, deliberately simplified to three parameters while acknowledging that a full operationalization would require five. Given this, the T3 criterion needs to address whether partial parameterization is always inconclusive or whether it can be a legitimate, acknowledged simplification. As written, the paper applies a stricter standard to Mireshghallah et al. than to Martin & Nissenbaum without explaining the difference, which weakens the force of that particular ✗ in Table 1.","section":"Section 3, T3: Five Parameters"}],"minor_comments":[{"comment":"The label 'same prompt sensitivity' is confusing; the phenomenon described is variation across repeated queries of the same prompt, so 'same-prompt variation' or 'sampling sensitivity' would be clearer.","section":"Section 4, 'Same Prompt Sensitivity'"},{"comment":"The sentence 'None of the prior works ... have accounted for prompt variation' lists only Mireshghallah et al., Ghalebikesabi et al., and Shao et al.; since the systematic claim appears to cover all nine surveyed works, the scope of this claim should be specified more precisely.","section":"Section 4, 'Paraphrasing Prompt Sensitivity'"},{"comment":"Figure 1 labels Level 1 as 'Interests & preferences of affected parties,' while the T4 text describes Level 1 as identifying 'winners and losers'; these descriptions should be aligned for consistency.","section":"Section 2, Figure 1 and T4 text"},{"comment":"In Section 4, 'Cao et al., 2024' is cited twice consecutively in the same parenthetical list; the duplicate citation should be removed.","section":"References"},{"comment":"The footnote states that the most accurate account of CI is the Privacy in Context book, but the tenets T1-T4 are drawn from later work (Nissenbaum 2019); the paper should clarify which source is authoritative for the four-tenet formulation.","section":"Footnote 1"}],"recommendation":"major_revision","confidential_remarks":"The paper is timely and the survey in Table 1 is a useful reference for the ML community. The main risk is that the categorical T2 and T4 claims will be seen as overreaching; a revision that reframes the standard as 'norm sources need validation and deliberative assessment' would preserve the central message and make it considerably more defensible. I have no concerns about novelty or citation fairness beyond the issues noted in the report."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis position paper does something genuinely useful: it states the four tenets of contextual integrity, audits nine LLM privacy papers against them in a clear table, and adds a prompt-sensitivity and position-bias checklist that most of those papers ignore. The audit itself is new, the paper is well-sourced, and it concedes that partial use of CI can still produce useful work, so it is not a strawman takedown.\n\nThe central claim is that existing LLM work \"inadequately applies\" CI, supported by a pattern of ✗ marks across T1-T4. That claim is too strong in one place. The load-bearing step is in Section 3, where the paper asserts \"Legal Statutes are not Privacy Norms\" and \"Crowdsourced Preferences are not Privacy Norms\" as categorical exclusions. That is contestable: law often codifies settled contextual norms (HIPAA is the obvious case), and surveys are the standard empirical method for eliciting norms in CI research — Martin & Nissenbaum 2016 is itself a survey instrument. The paper even acknowledges in Table 2 that governing institutions \"can be designed to support and reflect privacy norms,\" which sits in tension with the categorical T2 exclusion. If legal and preference data can serve as evidence of norms, then the blanket all-✗ T2 row and the inference from \"these works use proxies\" to \"they inadequately apply CI\" do not follow. The paper would be stronger if it argued that these sources are incomplete evidence that needs supplementation by normative analysis, not that they are never privacy norms.\n\nThe other soft spot is sample size: nine hand-selected works, mostly from 2023-24. That is reasonable for a position paper, but it cannot support a generalization to the whole LLM literature. Section 4's prompt-sensitivity discussion is useful, but the core empirical claim relies on the authors' own prior study; that is fine as an input, but it does not show that prompt variation would flip any audited conclusion.\n\nThis paper deserves a serious referee. It would benefit from a revision that softens the T2 categorical claims and explicitly allows partial CI use as legitimate borrowing. I would send it to reviewers, with a note to focus on the T2 tension.\n\nRecommendation: accept for peer review, expect revision.","headline":"A useful systematization of CI misuse in LLM papers, but the central charge overreaches by treating legal and crowdsourced proxies as categorically invalid.","tokens_in":15577,"tokens_out":1717,"would_cite":true,"duration_ms":16153,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM privacy evaluations are systematically under-applying Contextual Integrity, this paper argues.","keywords":["Contextual Integrity","LLM privacy","privacy norms","CI heuristic","prompt sensitivity","position bias","privacy evaluation","large language models"],"falsifier":"A controlled comparison could settle it: for a fixed set of LLM-agent information flows, compare privacy judgments obtained from the descriptive CI steps alone with judgments from a full CI heuristic run by an expert deliberative panel; if flows the descriptive steps flag as inappropriate are consistently judged legitimate by the heuristic, the paper's claim that skipping T4 yields flawed conclusions is confirmed, while full agreement would weaken it.","tokens_in":14665,"feed_emoji":"🔐","tokens_out":6381,"duration_ms":57452,"temperature":0.7,"pith_summary":"Contextual Integrity (CI) defines privacy as the appropriate flow of information, governed by privacy norms in a given social context. This position paper claims that nine recent studies using CI to evaluate large language models do not adhere to the theory's four core tenets: they substitute other privacy notions such as data minimization or secrecy, use legal statutes and crowdsourced preferences as proxies for privacy norms, sometimes omit the five CI parameters, and none apply the CI heuristic for normative legitimacy. The authors contend that partial or superficial use of CI can lead to incorrect privacy conclusions and flawed privacy-preserving system designs, and that current CI-based LLM experiments also overlook non-adversarial robustness issues such as prompt sensitivity and positional bias. A sympathetic reader takes away that CI-based privacy claims about LLMs are only as trustworthy as the theory's tenets they actually carry.","feed_headline":"LLM privacy studies misuse Contextual Integrity, authors argue","feed_subtitle":"Nine CI-based LLM evaluations skip the theory's normative tenets, risking invalid privacy claims.","key_machinery":"The evaluative machinery is the four-tenet checklist derived from CI theory, applied as a rubric to prior work. Each tenet does specific work: T1 fixes the object of analysis as flow appropriateness rather than data protection; T2 anchors appropriateness in contextual privacy norms rather than preferences or statutes; T3 makes the flow description complete and unambiguous; T4 supplies the normative step, the CI heuristic, that judges whether a norm-breaching flow is legitimate given affected interests, societal values, and contextual purposes. The paper's argument runs by marking each surveyed work against this rubric and showing that nearly all fail at least one tenet.","core_discovery":"On the authors' account, privacy under CI is not determined by data type, minimization, or legal compliance: T1 says privacy is appropriate information flow, T2 says appropriate flows conform with privacy norms, T3 says flows must be specified by all five parameters (sender, subject, recipient, information type, transmission principle), and T4 says the ethical legitimacy of a norm breach is assessed by the CI heuristic. Surveying nine works, the paper finds that eight deviate from T1, all nine deviate from T2, two deviate from T3, and all nine deviate from T4. Consequently, the paper concludes that these works borrow CI terminology without applying CI as a theory, making their privacy evaluations at best descriptive and at worst misleading for governance decisions.","pith_inferences":["An implication the paper leaves implicit is that the same under-application critique likely applies to CI-based privacy evaluations of systems beyond LLMs, such as IoT and data-sharing platforms, wherever legal rules or preference surveys stand in for norms.","If the position holds, legal statutes are not a safe shortcut for norms: privacy regulation and community norms can diverge, so a compliance-based privacy test could both over- and under-count violations.","A testable extension would be to annotate the same LLM information-flow scenarios with a full CI heuristic using expert deliberation and with the proxies criticized in the paper; the disagreement rate would quantify how much the under-application distorts conclusions.","The experimental-hygiene critique suggests that prompt and position sensitivity should be treated as first-class measurement axes in any LLM-based privacy study, not as noise to be averaged away."],"forward_implications":["Legal-compliance and preference-based LLM privacy evaluations should be described as measuring compliance or preference alignment, not privacy under CI, unless the normative step is carried out.","Future CI-based LLM benchmarks should require all five parameters and an explicit T4 evaluation; otherwise privacy-violation labels remain ambiguous.","LLM privacy measurements should report stability under paraphrased prompts and reordered answer options, since positional and prompt-induced biases can flip appropriateness judgments.","LLM agents that arbitrate information flows need a process for assessing the legitimacy of novel flows, not just a rule set, because a flow that breaches an existing norm may still be appropriate by contextual values.","Researchers who deliberately use only part of CI should state that their results are CI-inspired, avoiding the stronger claim that they evaluate privacy as defined by CI."],"supporting_citations":[{"why":"Foundational source of CI theory; supplies the definition of privacy as appropriate information flow and the account of social contexts.","marker":"Nissenbaum (2009)"},{"why":"States the four tenets T1-T4 and the assertion that privacy is not achieved through secrecy or data minimization.","marker":"Nissenbaum (2019)"},{"why":"Defines the CI heuristic used in tenet T4 for assessing the ethical legitimacy of norm-breaching flows.","marker":"Nissenbaum (2015)"},{"why":"Provides the survey methodology that prior CI-based LLM work draws on, and the caveat that full operationalization requires five-factor vignettes.","marker":"Martin & Nissenbaum (2016)"},{"why":"Grounds the claim that privacy norms are social goods, not individual preferences, used to reject crowdsourced preferences as norms.","marker":"Nissenbaum (2022)"},{"why":"Documents ambiguity in multidisciplinary interpretations of CI, supporting the paper's diagnosis of why the theory is under-applied.","marker":"Benthall et al. (2017)"},{"why":"Supplies the deliberative-panel idea the paper presents as a way to reach consensus on privacy norms for T2 and T4.","marker":"Susser & Bonotti (2024)"},{"why":"Reports measured variation in CI-based LLM responses under paraphrasing and Likert-scale changes, motivating the experimental-hygiene section.","marker":"Shvartzshnaider & Duddu (2025)"}],"fun_headline_variants":["LLM privacy studies misuse Contextual Integrity","Nine CI-based LLM evaluations skip the normative tenets","Contextual Integrity theory mishandled in 9 LLM papers","LLM privacy research borrows CI jargon without the theory","All nine LLM privacy studies violate CI's core tenets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The critique assumes that a genuine CI application must satisfy all four tenets with privacy norms kept separate from legal statutes and crowdsourced preferences; if partial use is a legitimate way to borrow the framework, the inadequate-application charge loses much of its force.","fun_headline_variants_meta":{"raw":{"variants":["LLM privacy studies misuse Contextual Integrity","Nine CI-based LLM evaluations skip the normative tenets","Contextual Integrity theory mishandled in 9 LLM papers","LLM privacy research borrows CI jargon without the theory","All nine LLM privacy studies violate CI's core tenets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1465,"prompt_tokens":824,"completion_tokens":641,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":562}},"tokens_in":440,"tokens_out":641,"duration_ms":6787,"temperature":1.0,"reasoning_tokens":562,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T21:01:55.355498+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled comparison could settle it: for a fixed set of LLM-agent information flows, compare privacy judgments obtained from the descriptive CI steps alone with judgments from a full CI heuristic run by an expert deliberative panel; if flows the descriptive steps flag as inappropriate are consistently judged legitimate by the heuristic, the paper's claim that skipping T4 yields flawed conclusions is confirmed, while full agreement would weaken it.","supporting_citations":[{"cited_title":"Privacy in context: Technology, policy, and the integrity of social life","cited_arxiv_id":null,"evidence_quote":"Foundational source of CI theory; supplies the definition of privacy as appropriate information flow and the account of social contexts."},{"cited_title":"Contextual integrity up and down the data food chain","cited_arxiv_id":null,"evidence_quote":"States the four tenets T1-T4 and the assertion that privacy is not achieved through secrecy or data minimization."},{"cited_title":"Respect for context as a benchmark for privacy online: What it is and isn't","cited_arxiv_id":null,"evidence_quote":"Defines the CI heuristic used in tenet T4 for assessing the ethical legitimacy of norm-breaching flows."},{"cited_title":"and Nissenbaum, H","cited_arxiv_id":null,"evidence_quote":"Provides the survey methodology that prior CI-based LLM work draws on, and the caveat that full operationalization requires five-factor vignettes."},{"cited_title":"Foreword by Helen Nissenbaum","cited_arxiv_id":null,"evidence_quote":"Grounds the claim that privacy norms are social goods, not individual preferences, used to reject crowdsourced preferences as norms."},{"cited_title":"Contextual integrity through the lens of computer science","cited_arxiv_id":null,"evidence_quote":"Documents ambiguity in multidisciplinary interpretations of CI, supporting the paper's diagnosis of why the theory is under-applied."},{"cited_title":"and Bonotti, M","cited_arxiv_id":null,"evidence_quote":"Supplies the deliberative-panel idea the paper presents as a way to reach consensus on privacy norms for T2 and T4."}],"review_version":1}