{"id":"c919c248-8f1a-4062-9e4f-5f3005d8f88b","arxiv_id":"2607.15084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Membership information in face-embedding geometry is controlled mainly by training-set size, not backbone or loss, and same-domain references show the signal shrinks as identity count grows.","lead":"Across 180 face recognition models, the number of training identities, not model size or loss function, dominates how much geometric cluster statistics reveal about whether a face was in the training set. The study adds a same-domain control showing cross-domain benchmarks inflate apparent membership leakage, which matters for privacy auditing of face recognition systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Same-domain reference is undefined at n_ids=all, yet Table 3 reports an 'all' AUC; the headline monotonic-decrease claim is not supported at the largest training-set size.","rationale":"The reader's conditional verdict already flags the same-domain split and the unexplained 'all' entry in Table 3. My stress test isolates the n_ids=all contradiction as the single most decisive issue: it is a concrete internal inconsistency rather than a vague concern about residual imbalance. If the all-level value is not from the same-domain reference, the abstract's monotonicity statement is overstated, although the 1K-to-100K trend and the n_ids-dominance result may still survive. This does not warrant a verdict change from CONDITIONAL; it strengthens the need for clarification and for restricting the claim to the measured levels. I marked agreement as partial because the reader's weakest assumption also emphasizes random-partition imbalances, while I focus on the undefined all-level reference as the primary load-bearing flaw.","tokens_in":17554,"tokens_out":11796,"duration_ms":133762,"concrete_test":"Request the per-model, per-benchmark AUC matrix (or recompute from the released code) and check whether any same-domain WebFace4M AUC exists for n_ids=all. If none exists, recompute Table 3's marginal means and the ANOVA excluding n_ids=all, and restrict the 'monotonic decrease' statement to the same-domain sequence n_ids=1K,10K,50K,100K. If the reported 'all' AUC of 0.71 was instead derived from cross-domain benchmarks, report that derivation explicitly and separate it from the same-domain analysis. The key pass/fail is whether the n_ids=all row can be reproduced from a same-domain non-member pool; if it cannot, the headline monotonic claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on the same-domain held-out WebFace4M split being a valid measurement at every n_ids level. Section 4.3 explicitly says this split samples non-members 'from the remaining subjects' and that 'at n_ids=all no non-members remain.' Nevertheless, Table 3 reports a marginal PairCos AUC of 0.71 for n_ids=all, and the abstract states the geometric membership signal 'decreases monotonically as more identities are added to training.' If the 'all' value was computed from cross-domain benchmarks rather than WebFace4M, then Table 3 mixes different non-member populations, and the monotonic sequence is not a same-domain curve. If it was computed some other way, the construction is unexplained and likely impossible. Either way, the same-domain evidence only establishes a decrease from 1K to 100K identities; the largest factor level either has no valid reference or is a cross-domain number. This is an internal inconsistency, not a disagreement with prior consensus. The Section 5.5 concession that random partitioning may correlate with identity-level attributes does not resolve the missing non-member pool at n_ids=all. Since the abstract's headline is precisely about same-domain monotonicity, this gap is load-bearing: the strongest form of the central claim is not supported by the presented measurement.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper quantifies how much membership information about training identities is retained in the geometry of face-recognition embedding clusters. The authors train 180 face-recognition models in a fully crossed design over three IResNet backbones, three angular-margin loss heads, five training-set sizes (1K/10K/50K/100K/all WebFace4M identities), and four training durations. For each model they compute four per-identity cluster statistics (pairwise cosine similarity, vMF concentration, penalized logit, and prototype softmax CE) and evaluate member/non-member separability (AUC) on nine benchmarks, including a same-domain held-out WebFace4M split and eight cross-domain benchmarks. The central claims are: (1) the number of training identities has the largest effect on separability, with backbone and loss head contributing little; (2) on the same-domain reference, the membership signal decreases monotonically with n_ids; (3) cross-domain benchmarks inflate the apparent signal; and (4) fusing all four statistics with a learned classifier recovers additional membership information. The paper also reports verification EERs to position the models within a realistic quality range.","tokens_in":17828,"tokens_out":5291,"duration_ms":55806,"significance":"If the claims hold, this is a valuable and unusually large-scale measurement contribution to the privacy-auditing literature. The 180-model factorial design, the matched-pair cross-domain comparisons (CFP-FF vs. CFP-FP, LFW vs. XQLFW), and the 5-fold stratified fusion evaluation are concrete strengths. The paper's main value is in identifying training-set size and training duration as the dominant controls on geometric membership leakage and in demonstrating that cross-domain benchmarks conflate domain shift with training-attributable signal. These are actionable findings for practitioners designing privacy audits. The paper is also honest about its limitations (single training source, no defenses, statistical assumptions), which increases confidence in its overall framing.","major_comments":[{"comment":"The same-domain held-out WebFace4M split is undefined at n_ids=all: Section 4.3 states 'at n_ids=all no non-members remain.' Yet Table 3 reports a marginal PairCos AUC of 0.71 for n_ids=all, and the abstract claims the geometric membership signal 'decreases monotonically as more identities are added to training' on the same-domain reference. This is internally inconsistent. The same-domain curve has no valid point at n_ids=all, so the monotonic decrease is only supported from 1K to 100K. If the 'all' value was computed from cross-domain non-members, then Table 3 mixes different non-member populations and the monotonic sequence is not a same-domain curve. The authors must either exclude n_ids=all from same-domain analyses (and revise the abstract/conclusion accordingly) or explicitly define and justify a same-domain non-member pool for n_ids=all. This is load-bearing for the headline clai","section":"Section 4.3, Table 3"},{"comment":"The abstract says 'the number of training identities has the largest effect on member/non-member separability,' but Table 3 shows that training duration (Epoch) has a comparable partial eta-squared: for PairCos both n_ids and Epoch are 0.99, and for PenLogit n_ids is 0.99 vs. 0.98 for Epoch. The paper's own text in Section 5.2 acknowledges that training duration is a close second. Without confidence intervals or a formal statistical comparison on the eta-squared values, the claim of 'largest effect' is not well supported. Please qualify the statement to reflect that n_ids and training duration both have dominant effects, with n_ids slightly larger, or provide uncertainty quantification showing the difference is significant.","section":"Section 5.2, Table 3, Abstract"}],"minor_comments":[{"comment":"The notation for the pairwise cosine statistic appears as both 'PairCos' (Table 3) and 'PAIRCOS' (text and Figure 2). Please standardize.","section":"Table 3 and throughout"},{"comment":"There are typographical artifacts such as 'ANOV A' and 'IRESNET' in tables; these should be fixed.","section":"General"},{"comment":"In the vMF concentration formula, \\bar{R} is used before it is defined; reorder the definitions for readability.","section":"Section 3.2"},{"comment":"The dashed 'Within-dist' curve is difficult to distinguish in grayscale; consider a different line style or a separate legend callout.","section":"Figure 2"},{"comment":"The probit-transform translation of Cohen's d' to AUC change is presented without a worked example; a short numerical illustration would help readers interpret the +0.26 and +0.22 values.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong fit for a security/privacy venue and the 180-model grid is a significant resource. The n_ids=all same-domain inconsistency is the primary blocker; it directly affects the abstract's monotonicity claim. The 'largest effect' wording regarding n_ids versus training duration also needs revision. These are fixable with revised claims and possibly additional analysis, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about this paper. First, it is a genuinely useful measurement study: 180 face recognition models trained in a fully crossed factorial design (3 backbones x 3 losses x 5 training sizes x 4 epoch counts), four per-identity cluster statistics, nine benchmarks, and a same-domain held-out WebFace4M split meant to isolate training-attributable leakage from domain shift. That scale and the same-domain control are the real contributions. Second, the paper's headline claim — that geometric membership signal decreases monotonically as more identities are added to training, on a same-domain reference — is not actually supported at the largest training-set size. Section 4.3 says at n_ids=all, no non-members remain in the held-out split. Yet Table 3 reports an 'all' PairCos AUC of 0.71. Where did those non-members come from? The paper never says. If they came from cross-domain benchmarks, the same-domain curve breaks. If they came from somewhere else, the construction is missing. Either way, the monotonic decrease is only established from 1K to 100K. The 100K value is 0.71 and 'all' is also 0.71, so the curve may just be flat, but the abstract overreaches.\n\nWhat the paper does well: the factorial design is the largest systematic study in embedding-level MIA; the matched-pair cross-domain comparisons (CFP-FF vs CFP-FP, LFW vs XQLFW) cleanly quantify how pose and quality inflate apparent leakage; and the authors are honest that PAIRCOS is equivalent to Li et al.'s intra-similarity. The core practical finding — training-set size dominates, backbone and loss contribute little — is visible in the tables and consistent across all 180 models. I think that finding holds up.\n\nSoft spots, in order: (1) the n_ids=all gap is a real internal inconsistency, load-bearing for the abstract. (2) All AUCs are point estimates with no confidence intervals; given 180 models, error bars are easily obtainable. (3) The ANOVA response is underspecified — what exactly is the unit, a single model's AUC per benchmark? The paper gives partial eta-squared but not how the response was formed. (4) No code, model checkpoints, or precomputed statistics released, which limits the 'measurement' value.\n\nThe paper is for privacy auditors and people working on membership inference for embedding models. It deserves a serious referee, not a desk reject. I'd send it to review with the request that the authors fix the n_ids=all reporting, add uncertainty quantification, and make the abstract match the evidence.","headline":"Useful large-scale measurement study, but the same-domain monotonic claim overreaches at n_ids=all.","tokens_in":18312,"tokens_out":2893,"would_cite":true,"duration_ms":30945,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A face recognition model's embedding geometry carries measurable membership information, and the amount is controlled mainly by the number of training identities rather than by backbone or loss head.","keywords":["membership inference","face recognition","embedding geometry","hypersphere","privacy auditing","angular-margin loss","cluster statistics","training-set size"],"falsifier":"Train a batch of models at 1K and 100K identities on a different large face dataset and check whether the same-domain held-out gap between member and non-member AUC still decreases monotonically; a non-monotonic or inverted trend would refute the claim that identity count is the dominant control. Alternatively, construct a held-out split that matches image count and quality per identity and see if the separation vanishes or changes.","tokens_in":17403,"feed_emoji":"🕵️","tokens_out":3395,"duration_ms":33846,"temperature":0.7,"pith_summary":"This paper asks how much a face recognition model's output embeddings reveal whether a given identity was part of its training data, when only the embeddings are available. It trains 180 models varying backbone size, loss head, training duration, and the number of training identities, and measures four cluster-geometry statistics that separate member from non-member identities. The dominant finding is that the number of training identities controls the separation far more than architecture or loss: on a same-domain held-out reference, the membership signal falls monotonically as more identities are added. Cross-domain non-member sets (different pose, quality, age, ethnicity) inflate the apparent signal, so privacy audits need same-domain references. Combining all four statistics with a learned classifier recovers membership information beyond any single statistic.","feed_headline":"More training identities shrink face-recognition privacy leaks","feed_subtitle":"A 180-model study finds face-embedding geometry reveals training membership, and cross-domain benchmarks exaggerate the leak.","key_machinery":"The central object is the per-identity embedding cluster on the unit hypersphere, summarised by four statistics: pairwise cosine similarity (cluster tightness), von Mises-Fisher concentration kappa, a leave-one-out margin-penalised logit approximating the training loss, and a prototype softmax cross-entropy treating other identities as negatives. The argument runs on the assumption that angular-margin losses optimise only training identities, leaving non-member clusters with looser geometry; the statistics convert that geometry into scalar scores, and threshold-classifier AUC quantifies how much membership information survives. The factorial design over 180 models isolates which training-tim","core_discovery":"The paper establishes that the hyperspherical embedding clusters of training identities remain statistically distinguishable from those of unseen identities, and that the size of this distinguishability is governed chiefly by the number of training identities rather than by backbone or loss head. Using four per-identity statistics—mean pairwise cosine similarity, a von Mises-Fisher concentration estimate, a margin-penalised logit, and a prototype softmax cross-entropy—the authors measure threshold-classifier AUC separating members from non-members across 180 models. On a same-domain held-out split, AUC decreases monotonically as training identities grow from 1K to all available identities, w","pith_inferences":["The monotonic decrease with training-set size suggests a scaling law: leakage may continue to drop with even larger identity counts, potentially below practical attack thresholds, though the paper does not test that.","The finding that cross-domain benchmarks inflate leakage implies that prior membership-inference studies on face recognition may have overstated real-world risk for models trained on large public data; replicating those attacks with same-domain references could change reported accuracy numbers.","Because the statistics are cluster-level and require only a few probe images per identity, an auditor could probe any deployed model with a handful of photos; the fusion result suggests a practical auditing tool, but also means the same technique could be used adversarially.","A testable extension would train models on different data sources (the paper itself flags this as open) to see whether the monotonic trend and factor ranking hold across data distributions."],"forward_implications":["Privacy auditors should use same-domain held-out non-members rather than cross-domain benchmarks, since cross-domain sets overstate the membership leak.","Training on more identities is the only tested design choice that substantially reduces geometric membership leakage; backbone and loss head have minor effects.","Even at large training-set sizes where individual statistics are weak, combining geometry statistics with a classifier reveals additional membership information, so audits should fuse statistics.","The measured separability is a lower bound on membership information available from embeddings; richer models could extract more.","The monotonic decrease implies that models trained on small identity pools (e.g., niche or custom datasets) are the most vulnerable to this type of geometric membership inference."],"fun_headline_variants":["Face-embedding geometry separates training from unseen identities","Training identity count, not model size, sets face-recognition leak","Cross-domain benchmarks overstate face-recognition membership leaks","More training identities shrink face-recognition membership signal","Fused geometry stats extract hidden face-membership information"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The same-domain held-out reference is built by randomly splitting one dataset into member and non-member identities, assuming the random split removes all domain shift; if the split leaves residual differences in image count, quality, or pose, the monotonic decrease and the cross-domain inflation conclusions are confounded.","fun_headline_variants_meta":{"raw":{"variants":["Face-embedding geometry separates training from unseen identities","Training identity count, not model size, sets face-recognition leak","Cross-domain benchmarks overstate face-recognition membership leaks","More training identities shrink face-recognition membership signal","Fused geometry stats extract hidden face-membership information"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000272,"raw_usage":{"total_tokens":1457,"prompt_tokens":718,"completion_tokens":739,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":660}},"tokens_in":462,"tokens_out":739,"duration_ms":8642,"temperature":1.0,"reasoning_tokens":660,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:11:57.393224+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a batch of models at 1K and 100K identities on a different large face dataset and check whether the same-domain held-out gap between member and non-member AUC still decreases monotonically; a non-monotonic or inverted trend would refute the claim that identity count is the dominant control. Alternatively, construct a held-out split that matches image count and quality per identity and see if the separation vanishes or changes.","supporting_citations":[],"review_version":1}