{"id":"29fb5857-0dae-4b75-8612-dbe3846255a5","arxiv_id":"2607.05318","paper_version":1,"verdict":"CONDITIONAL","confidence":"UNKNOWN","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A benchmark of 85 manually curated workplace scenarios reveals that multi-user AI agent systems suffer high rates of contextual integrity violations across outputs, inter-agent communication, and shared memory.","lead":"PiSAs is a new benchmark for measuring unintentional privacy leaks in multi-user AI agent systems, where shared infrastructure (memory, inter-agent messages) can expose sensitive information across users. It reveals that even state-of-the-art LLMs fail to reliably filter inappropriate content, and system design mitigations often just relocate leaks rather than eliminate them.","discovery_kind":"new_method","skeptic_critique":{"model":"glm-5.2","headline":"Violation rates depend on an LLM-judge pipeline whose lenient extraction, partial-match verification, and any-K union aggregation could systematically inflate the central claim that no configuration achieves acceptable privacy preservation.","rationale":"The reader correctly identified the LLM-judge pipeline as the load-bearing concern. My analysis confirms this and sharpens the mechanism: the interaction between lenient partial-match verification and any-K union aggregation creates systematic upward pressure on absolute violation rates, and the paper's own Appendix I shows nearly 2x shifts when verifier ensembles change. However, I recommend UNCHANGED rather than a harsher verdict for three reasons. First, the paper itself acknowledges sensitivity (Appendix I) and frames absolute values cautiously, emphasizing qualitative trends. Second, the worked examples (Appendix J) demonstrate genuine violations that are clearly inappropriate, confirming the phenomenon is real even if rates are uncertain. Third, the attribute-type analysis (Figure 4a) shows a sensible pattern — task-entangled attributes leak far more than irrelevant ones — which would be unlikely under pure measurement noise. The CONDITIONAL verdict is appropriate: the qualitative finding (violations are pervasive and mitigations relocate rather than eliminate them) is robust, but the specific claim that 'no configuration achieves acceptable privacy preservation' depends on absolute rates that are not yet stable enough to fully endorse without human validation. The secondary concern about LLM-generated visibility labels for Severity Classification scenarios adds modest additional uncertainty but affects only 25/85 scenarios and only the visibility metric.","tokens_in":44165,"tokens_out":3395,"duration_ms":120551,"concrete_test":"Sample 100 flagged violations stratified across configurations (single/centralized/decentralized × 3 backbones) from the main results. Have 3 independent human annotators, given the same ground-truth attributes and visibility matrices, label each flagged violation as genuine or false positive. If the false positive rate exceeds 15%, recompute the any-K aggregated violation rates with corrected counts; if the failure rate (Table 2, runs with ≥1 violation) drops below 50% for the best configuration (centralized + Claude), the claim that 'no configuration achieves acceptable privacy preservation' would need qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that no studied configuration achieves acceptable privacy preservation — rests on violation rates produced by a three-stage pipeline (Appendix B): (1) a lenient extractor flags candidate attribute mentions without ground truth, (2) three verifier models with ground-truth access vote by majority on whether each candidate 'reveals the key fact, even partially,' and (3) violations are aggregated via any-K union across K=3 runs (Appendix G). Each stage introduces upward pressure on reported rates. The 'even partially' criterion sets a low bar: an agent that says 'Eli has a medical condition affecting room choice' might be flagged as partially revealing the raw attribute 'documented hyperacusis diagnosis,' even though this is closer to the sanitized operational form. The any-K union then compounds any per-run false positives: with a per-run false positive rate of just 0.15 per attribute, the 3-run union yields ~38% apparent violation for that attribute. The paper's own sensitivity analysis (Appendix I, Tables 24–25) confirms substantial sensitivity: replacing the paid verifier ensemble with local models shifts centralized V_vis from 22.9% to 42.9% and V_appr from 26.2% to 42.9% — nearly doubling. While qualitative trends (Claude > other backbones, centralized < decentralized for visibility) survive across stacks, the absolute rates that underpin 'no acceptable configuration' are not stable. The worked example in Appendix J does show genuine violations (e.g., Eli's hyperacusis diagnosis propagated verbatim), so violations are real — but the question is whether the reported rates are inflated enough to make the headline claim stronger than the evidence supports. A secondary concern: for 25 of 85 scenarios (Severity Classification), visibility ground-truth labels are LLM-generated (Qwen3.6-27B, Appendix A.3), meaning visibility violations are measured against a potentially noisy reference.","agreement_with_reader":"agree"},"referee_report":{"model":"glm-5.2","summary":"The paper introduces PiSAs, a benchmark for evaluating contextual integrity (CI) violations in multi-user agentic systems where multiple users share the same infrastructure (agents, memory, communication channels). Unlike prior multi-user CI benchmarks that study independently owned agents negotiating via dialogue, PiSAs evaluates spillage through shared infrastructure. The benchmark comprises 85 manually curated workplace scenarios across three task families (JIRA allocation, meeting allocation, severity classification), each annotated with dual CI labels: task-contextual appropriateness and per-person visibility. The authors evaluate three agent topologies (single, centralized, decentralized), three memory configurations, three LLM backbones, and varying privacy-prompt levels. The central finding is that no studied configuration achieves acceptable privacy preservation: appropriateness violations remain pervasive (e.g., V_appr exceeding 77% for single-agent systems), partitioning reduces but does not eliminate leakage, and memory-based mitigations relocate violations to memory stores rather than eliminating them. The paper also finds that task-entangled personal attributes are the hardest to handle, and that strict rule specification outperforms broad policy guidance.","tokens_in":44387,"tokens_out":1819,"duration_ms":179009,"significance":"The paper addresses a genuine gap in the CI benchmarking literature: prior work either studies single-user settings or independently owned multi-agent systems, but does not consider the shared-infrastructure multi-user setting where spillage occurs through memory and inter-agent communication. The dual-annotation scheme (appropriateness + visibility) is well-motivated and enables measurement of qualitatively distinct failure modes. The systematic experimental design—varying topology, memory, backbone, and prompt defenses—is a strength, as is the transparent reporting of the LLM-judge pipeline and its sensitivity. The worked example (Appendix J) concretely demonstrates genuine violations, lending credibility to the evaluation. The finding that memory relocates rather than eliminates violations is a useful and actionable insight for system designers. The benchmark's system-agnostic design and the commitment to release data under Apache 2.0 are additional strengths.","major_comments":[{"comment":"§5, Table 2, and Appendix I, Tables 24–25: The central claim that 'no studied configuration achieves an acceptable level of privacy preservation' (§6) rests on absolute violation rates produced by the LLM-judge pipeline. The paper's own sensitivity analysis shows that replacing the verifier ensemble nearly doubles centralized V_vis from 22.9% to 42.9% and V_appr from 26.2% to 42.9% (Table 24, centralized column). While the authors state that 'qualitative trends remain consistent,' the claim of 'no acceptable configuration' is a statement about absolute levels, not just trends. The paper does not define what 'acceptable' means quantitatively, making the central claim unfalsifiable as stated. The authors should either (a) define an explicit threshold for 'acceptable' privacy preservation and show it is exceeded under all judge configurations, or (b) reframe the conclusion to acknowledge it","section":null},{"comment":"§5 and Appendix B: The 'even partially' verification criterion (Appendix B: verifiers vote on whether each candidate 'reveals the key fact, even partially') sets a low bar that could conflate sanitized operational abstractions with raw attribute leaks. For example, an agent stating 'Eli has a medical condition affecting room choice' might be flagged as partially revealing 'documented hyperacusis diagnosis,' even though this is closer to the sanitized form ('Meetings requiring Eli should not use Delta'). The paper acknowledges sanitized forms in Appendix A and the worked example (Appendix J) shows the centralized system producing 'Eli has a medical condition making the Delta room unsuitable' — which appears to be a reasonable sanitization, yet would likely be flagged. The paper should report what fraction of flagged violations are partial/paraphrased leaks versus near-verbatim raw-attr","section":null},{"comment":"§5.2, Figure 3: The claim that memory 'relocates violations rather than eliminating them' is supported by the observation that total visibility violations rise from 36–47% to 63–90% with Hybrid memory. However, the paper also notes that the reduction in A2A violations is 'confounded with a reduction in the number of A2A messages' (agents read from memory instead of communicating). The paper does not report whether the total number of attribute-exposure opportunities is comparable across memory conditions. If memory reduces total information flow volume, a fair comparison would normalize violations by exposure opportunities (e.g., violations per attribute-transmission event). Without this, it is unclear whether memory genuinely increases total violations or whether the higher count reflects more surfaces being audited. The authors should either add this normalization or qualify the claim.","section":null}],"minor_comments":[{"comment":"Abstract: 'PiSAsis' — missing space before 'is'.","section":null},{"comment":"Abstract: 'We introducePiSAs' — missing space before 'PiSAs'.","section":null},{"comment":"Table 1: The 'Acc.' column uses 'T' and 'U' for accumulation type, but the legend only explains these in the caption indirectly ('cross-task T, cross-user U or – not measured'). A reader unfamiliar with the notation may find this unclear on first reading.","section":null},{"comment":"§3, Scenario Construction: The paper states '85 scenarios' but Table 3 reports 25+25+35=85. This is consistent, but the distribution across task families is uneven (Severity has only 3 appropriate attributes per scenario). The paper should discuss whether this imbalance affects the aggregated metrics reported in Table 2.","section":null},{"comment":"Appendix A.3: Visibility annotations for Severity Classification are assigned using Qwen3.6-27B as an LLM judge, while JIRA and Meeting allocations use manual annotation. This inconsistency should be noted in the main text (§3), not only in the appendix, as it affects the reliability of visibility violation rates for the Severity task.","section":null},{"comment":"Table 2: The 'Failure rate' column is defined as 'fraction of runs with at least one appropriateness violation' but the caption does not make clear whether this is per-run or per-scenario (union across K=3). Appendix D.4 clarifies it is per-run, but the main table caption should state this.","section":null},{"comment":"Figure 3: The y-axis label 'Violation (%)' is ambiguous — it is unclear whether this is V_appr, V_vis, or both. The figure caption should specify which metric is plotted for each bar group.","section":null},{"comment":"Appendix G, Tables 19–20: The any-K aggregation is justified well, but the paper could note that with K=3, the any-K union inflates rates relative to mean-run by approximately 1.3–1.5× (based on the tables). This quantitative characterization would help readers calibrate the absolute numbers.","section":null},{"comment":"References: Several entries cite 2026 papers (e.g., Anthropic 2026, Park et al. 2026, Fu et al. 2026). If these are accepted/published, please verify venue and page numbers; if preprints, please mark as 'Preprint' consistently.","section":null},{"comment":"§4.1, Single configuration: The paper states 'we simulate this system by giving an LLM unrestricted access to all available information in its context (except the ground-truth labels).' It would help to clarify whether the single agent receives all attributes in a single prompt or through simulated multi-turn interaction.","section":null}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about the LLM-judge pipeline inflating violation rates is partially valid: the sensitivity analysis in Appendix I does show substantial absolute metric variation, and the 'even partially' criterion is genuinely lenient. However, the worked example in Appendix J demonstrates real violations (e.g., the single-agent system propagating 'documented hyperacusis diagnosis' verbatim), so the concern does not undermine the paper's qualitative findings. The main issue is that the central claim ('no acceptable configuration') is stated as an absolute without a defined threshold, while the supporting evidence shows the absolute numbers are judge-dependent. This is fixable by reframing the conclusion or adding a threshold analysis. The paper is a solid contribution to an underexplored area (shared-infrastructure multi-user CI) and should be publishable after revision."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for a careful and constructive report. The referee correctly identifies three substantive issues: (1) the central 'no acceptable configuration' claim is not tied to a defined threshold and is sensitive to judge choice; (2) the 'even partially' verification criterion may conflate sanitized abstractions with raw attribute leaks; and (3) the memory 'relocation' claim is not normalized by exposure opportunities. We agree with all three points and will revise the manuscript accordingly. Specifically, we will define an explicit acceptability threshold, add a breakdown of partial vs. near-verbatim leaks, and add a normalized violation-rate analysis for the memory comparison. No standing objections remain.","responses":[{"response":"The referee is correct on both points. The term 'acceptable' is used informally and is not tied to a defined threshold, which makes the central claim unfalsifiable as stated. We also agree that the sensitivity analysis in Appendix I shows substantial absolute-level variation under different judge/verifier configurations, which complicates any claim resting on absolute rates. We will make two changes. First, we will define an explicit acceptability threshold. We propose adopting a per-attribute violation rate threshold of 5% (i.e., at most 5% of inappropriate attributes leaked, and at most 5% of hidden attributes exposed to unauthorized users), motivated by the severity of privacy harm from even single-attribute leaks in organizational settings. Under this threshold, even the most conservative judge configuration (Paid J + Paid V) yields V_appr of 25.4% for the best centralized configuration (Table 2), far exceeding 5%. The alternative judge configuration that produces the lowest centralized V_appr (Gemma J + Paid V: 22.8%) still exceeds the threshold by a factor of 4.5. No configuration approaches this bar under any judge configuration we tested. Second, we will reframe the conclusion to explicitly state the threshold and acknowledge the judge sensitivity: 'Under a 5% per-attribute violation threshold, no studied configuration achieves acceptable privacy preservation under any judge configuration tested, though absolute violation rates vary substantially across evaluation stacks.' We believe this makes the claim both falsifiable and appropriately qualified.","revision_made":"yes","referee_comment":"The central claim that 'no studied configuration achieves an acceptable level of privacy preservation' rests on absolute violation rates from the LLM-judge pipeline, but the sensitivity analysis shows judge replacement nearly doubles some metrics, and 'acceptable' is undefined, making the claim unfalsifiable."},{"response":"We agree this is an important distinction and that the current presentation does not adequately separate these cases. The referee's example is well-taken: a statement like 'Eli has a medical condition making the Delta room unsuitable' (which appears in the centralized system's output in Appendix J) is arguably a reasonable sanitization, yet under our 'even partially' criterion it would likely be flagged as partially revealing the raw attribute 'documented hyperacusis diagnosis.' We will address this in two ways. First, we will add a breakdown classifying flagged violations into three categories: (a) near-verbatim raw-attribute leaks (the sensitive detail is propagated essentially as-is), (b) partial/paraphrased leaks (some identifying detail from the raw attribute is present, but not the full sensitive fact), and (c) sanitized-abstraction leaks (only the operational implication is conveyed, which our current pipeline may still flag). We will produce this breakdown by running a post-hoc categorization pass on the existing flagged violations. Second, we will discuss the limitation that our verification pipeline does not distinguish sanitized abstractions from partial leaks, and note that this may inflate violation rates for task-entangled attributes where the system produces a reasonable sanitization. We note that even after accounting for this, the violation rates for clearly irrelevant personal attributes (V_appr = 19.3% averaged over systems, Table 12) — where sanitization is not applicable — remain well above any reasonable acceptability threshold, so the core finding is not solely an artifact of the partial-leak criterion.","revision_made":"yes","referee_comment":"The 'even partially' verification criterion sets a low bar that could conflate sanitized operational abstractions with raw attribute leaks. The paper should report what fraction of flagged violations are partial/paraphrased leaks versus near-verbatim raw-attribute leaks."},{"response":"This is a fair concern. The current comparison is between raw violation counts (or rates) across memory conditions that differ in the number of information-transmission events. As the referee notes, Hybrid memory reduces A2A messages from ~11 to ~4 (Table 10), so comparing raw A2A violation rates across conditions is not apples-to-apples. We will add a normalized analysis. Specifically, we will compute violations per attribute-exposure opportunity, where an exposure opportunity is defined as an instance where an attribute could potentially surface at a given surface (e.g., one A2A message, one memory write event). This normalizes for the fact that memory changes both the volume and the location of information flow. We will report this alongside the raw rates. We expect that the normalized analysis will show that memory reduces per-opportunity violation rates on the A2A channel (consistent with the current observation that A2A violations drop sharply) but that the memory channel itself has a high per-opportunity violation rate, supporting the 'relocation' claim in a more controlled way. However, we will also qualify the claim if the normalized analysis shows that the total per-opportunity violation rate does not increase with memory. We agree that without this normalization, the strong form of the 'relocation' claim is not fully supported.","revision_made":"yes","referee_comment":"The claim that memory 'relocates violations rather than eliminating them' is confounded by a reduction in A2A message volume. The paper does not normalize violations by exposure opportunities, making it unclear whether memory genuinely increases total violations or whether the higher count reflects more surfaces being audited."}],"tokens_in":44524,"tokens_out":1259,"duration_ms":167156,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Here's my read on PiSAs. The paper introduces a genuinely new evaluation setting: multiple users sharing the same agentic infrastructure, where privacy violations can occur not just in outputs but through shared memory and inter-agent communication. Prior multi-user CI work (MAGPIE, PAC-Bench) studied independently owned agents negotiating; this paper studies shared infrastructure where spillage happens through the plumbing. That distinction matters and the benchmark design reflects it well — dual annotations for appropriateness (should this attribute be used for this task?) and visibility (should this user see this attribute?) are the right axes, and the scenario construction with oracle solutions is careful work. The finding that memory and partitioning relocate rather than eliminate violations is the kind of result practitioners need to hear. The worked example in Appendix J is genuinely illustrative — you can see Eli's hyperacusis diagnosis propagating verbatim through agent messages. The ablations on rule explicitness and attribute cueing are useful diagnostics. The stress-test concern about the LLM-judge pipeline is real but partially overblown. The any-K union aggregation does inflate absolute rates relative to mean-run, and the sensitivity analysis (Appendix I) shows that swapping to local verifiers nearly doubles some metrics. But the paper is transparent about this — they report the sensitivity analysis, they show that qualitative rankings (Claude > others, centralized < decentralized for visibility) survive across judge stacks, and the any-K choice is defensible under a privacy interpretation where any demonstrated leakage path counts. The concern about the 'even partially' verification criterion is worth raising but the worked example shows real verbatim leaks, not borderline paraphrases. Where I do think the headline overstates: 'no configuration achieves acceptable privacy preservation' is stronger than the evidence supports, because the absolute rates are not stable across evaluation stacks and 85 scenarios across 3 task types is modest. The paper would benefit from framing this as 'violation rates remain high across all studied configurations' rather than a categorical claim. The LLM-generated visibility labels for the Severity Classification task (Appendix A.3) are a minor concern — 25 of 85 scenarios use Qwen-generated visibility annotations, which means visibility violations for those scenarios are measured against a potentially noisy reference. Overall: solid benchmark, careful experimental design, honest about limitations. The absolute numbers should be read as ordinal rather than cardinal. Worth a serious referee — the benchmark fills a real gap and the findings are directionally important for anyone building multi-user agent systems.","headline":"New benchmark for cross-user privacy spillage in shared agentic systems; absolute violation rates are judge-sensitive but qualitative findings hold.","tokens_in":45222,"tokens_out":571,"would_cite":true,"duration_ms":150940,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"No multi-user agent system preserves privacy under shared use","keywords":[],"falsifier":"If an LLM-judge configuration were found that systematically reversed the relative ranking of system designs or backbones on violation rates, or if a deployed system with the same architecture achieved near-zero violations on these scenarios through a mitigation not studied in the paper, the central claim that no configuration achieves acceptable privacy would be weakened.","tokens_in":44480,"feed_emoji":"🔐","tokens_out":960,"duration_ms":131991,"temperature":0.7,"pith_summary":"This paper introduces PiSAs, a benchmark for measuring unintentional privacy leakage when LLM agents serve multiple users who share the same agentic infrastructure. The central insight is that existing privacy benchmarks, grounded in contextual integrity theory, focus on single-user settings or independently owned agents negotiating with each other, and therefore miss a qualitatively new risk: when users share agents and memory, sensitive information spills across users through the system's own internal channels, not just through outputs to external recipients. PiSAs addresses this with dual annotations on every piece of information: whether it is appropriate to use for the current task, and which users are legitimately allowed to see it. This lets the benchmark track two distinct failure modes: appropriateness violations, where the system uses information it should not, and visibility violations, where information reaches users who should not have access to it. The authors evaluate three agent topologies (single, centralized, decentralized), three memory configurations (none, private, shared/hybrid), and three LLM backbones, and find that no configuration achieves acceptable privacy preservation. Data and agent partitioning reduces leakage but does not eliminate it, and adding memory shifts violations from communication channels into memory stores rather than removing them. The paper identifies the core bottleneck as the LLM's own judgment: even the strongest models fail to reliably filter inappropriate content or restrict information to authorized users, particularly when sensitive attributes are entangled with task completion.","feed_headline":"No multi-user agent system preserves privacy under shared use","feed_subtitle":"A new benchmark shows sensitive data spills across users through agent messages and shared memory in every system design tested, with LLMs, ","key_machinery":"The benchmark's machinery is the dual annotation scheme: each attribute in a scenario carries a task-contextual appropriateness label (whether it should be used for the task) and a per-person visibility structure (which users may legitimately access it). These two orthogonal annotations, combined with evaluation at three leakage surfaces (agent-to-agent communication, memory stores, and gathered-information summaries), allow the benchmark to distinguish between appropriateness violations and visibility violations, and to detect when mitigations relocate rather than eliminate leaks. The scenarios are constructed backward from a hidden oracle solution, ensuring that a privacy-preserving answer","core_discovery":"The central discovery is that in multi-user agentic systems, privacy violations are pervasive and structural rather than incidental. Partitioning information across agents and restricting communication topology reduces but does not eliminate violations, because the LLM's judgment calls remain the bottleneck. When memory is introduced to improve task performance, violations migrate from agent-to-agent communication channels into memory stores, making memory a persistent and harder-to-defend leakage surface. The hardest case for systems is not irrelevant personal information but decision-critical personal attributes that must be used in sanitized form yet are frequently propagated in their raw","pith_inferences":[],"forward_implications":["Multi-user agentic systems deployed in organizational settings today likely already exhibit the cross-user spillage this benchmark measures, since no studied configuration achieves acceptable privacy preservation.","Memory systems in multi-agent architectures function as unregulated data lakes for sensitive information, and privacy-preserving memory designs that enforce visibility constraints at the storage layer rather than relying on LLM judgment are needed.","The finding that broad policy guidance performs worse than no guidance at all suggests that naturalistic policy documents may license information use rather than constrain it, with implications for how organizations instruct both human and AI agents.","Fine-tuning on contextual integrity reasoning, rather than prompt-level instructions alone, may be necessary to close the judgment gap the paper identifies as the primary bottleneck.","The benchmark's design, which separates appropriateness from visibility, could be extended to cross-task leakage scenarios where information legitimately accessed for one task surfaces inappropriately in a later one."],"fun_headline_variants":["Shared memory makes agent privacy leaks harder to catch","Partitioning agents cuts some privacy leaks but LLM judgment stays the bottleneck","Multi-user agents leak sensitive data across every topology tested","Agent memory becomes a persistent leakage surface for personal data","State-of-the-art LLMs fail to filter inappropriate content across users"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The evaluation pipeline relies on LLM judges to detect whether privacy violations occurred. If these judges systematically over-count or under-count violations, for example by hallucinating matches or missing paraphrased leaks, the reported violation rates could be artifacts of the evaluation method rather than properties of the agentic systems being tested. The authors note that absolute metric values vary across judge configurations, though they claim the main findings hold","fun_headline_variants_meta":{"raw":{"variants":["Shared memory makes agent privacy leaks harder to catch","Partitioning agents cuts some privacy leaks but LLM judgment stays the bottleneck","Multi-user agents leak sensitive data across every topology tested","Agent memory becomes a persistent leakage surface for personal data","State-of-the-art LLMs fail to filter inappropriate content across users","Restricting agent communication reduces but cannot eliminate cross-user spillage","Adding memory shifts agent privacy leaks from messages into stored data","Decision-critical personal attributes are the hardest case for agent privacy","No tested agent architecture prevents cross-user data spillage","Dual contextual-integrity annotations expose where agents leak user data"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":671,"prompt_tokens":504,"completion_tokens":167,"prompt_tokens_details":null},"tokens_in":504,"tokens_out":167,"duration_ms":11747,"temperature":1.0,"reasoning_tokens":48,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T17:40:21.769325+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If an LLM-judge configuration were found that systematically reversed the relative ranking of system designs or backbones on violation rates, or if a deployed system with the same architecture achieved near-zero violations on these scenarios through a mitigation not studied in the paper, the central claim that no configuration achieves acceptable privacy would be weakened.","supporting_citations":[],"review_version":1}