{"id":"370dbfbd-93c0-4af8-be93-89a1dc5a8b65","arxiv_id":"2506.05769","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Brain fingerprinting metrics that only compare average self-to-self and self-to-other similarity do not capture recognition accuracy; standard biometric measures like EER should be used instead.","lead":"Neuroscientists often call brain connectivity patterns 'fingerprints,' but many of the measures used do not actually test whether one person can be told apart from another. This paper argues that popular metrics like Idiff and Dself can look excellent even when a system would make many recognition errors, and recommends using standard biometric error rates instead.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified: the mathematical claim that Idiff and Dself ignore distribution overlap is sound, and the Beta simulation is a valid counterexample; absence of real data limits practical prevalence estimates but not the core argument.","rationale":"The reader identified the absence of real-data validation as the main weakness, making the verdict CONDITIONAL. I agree that real-data validation would strengthen the practical implications, but I do not see it as load-bearing for the central claim. The paper's argument is essentially a counterexample: Idiff and Dself are functions of mean differences, while recognition performance depends on distribution overlap, so the two can disagree. The Beta simulation in Section IV provides concrete instances, and the paper does not overclaim that all real datasets will show the discrepancy, only that it cannot be excluded. This logical point holds regardless of whether real connectome distributions are Beta-distributed, because the simulation is an existence proof rather than a population estimate. Section V's statement that 'based on the distributions of the real values, we will find ourselves in one of the possible cases outlined by the different distributions reported by the simulation' is acceptable: it does not assert that the problematic cases are guaranteed, only that they are possible. The practical recommendation to prefer rank-k or EER is also consistent with the argument, since those metrics directly operationalize recognition errors. The absence of real data is a limitation in scope, not a flaw in the reasoning. I would therefore keep the reader's CONDITIONAL verdict unchanged, while noting that a real-data test would be a useful, though not logically necessary, robustness check.","tokens_in":6088,"tokens_out":4377,"duration_ms":50928,"concrete_test":"On a public test-retest connectome dataset (e.g., HCP, 100 unrelated subjects, two sessions), compute functional-connectivity identifiability matrices for several standard atlases and preprocessing choices; for each, compute Idiff, mean Dself, and EER from the matching-score distributions. Then check whether the EER-vs-Idiff scatter contains cases with EER approximately 0 and low Idiff, or high EER and high Idiff. Finding such cases in real data would settle that the synthetic concern occurs empirically; finding none across many settings would show practical harm is rarer but would not refute the mathematical point.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that Idiff and Dself, as mean-separation statistics, cannot summarize recognition performance because overlap, dispersion, and shape of genuine/impostor score distributions matter. This is true by construction and is demonstrated by the 256 Beta scenarios in Section IV: EER approximately 0 pairs with both low and high Idiff, and high EER pairs with similar Idiff values. The claim is an existence/counterexample argument, so the lack of real connectome data (Section IV: 'Including an analysis based on real data seems superfluous in this context') is not load-bearing for the logical point; it only limits estimates of how often practical studies are misled. The paper's wording ('could lead', 'may provide') is appropriately modal. I therefore identify no significant objection to the central argument. The weakest point remains the practical generalization, and a real-data check would quantify it, but it does not undermine the paper's conclusion.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript argues that connectome 'brain fingerprinting' metrics such as differential identifiability (Idiff) and Dself, which are based on mean differences between within- and between-subject similarities, do not adequately characterize biometric recognition performance because they ignore the overlap, dispersion, and shape of genuine and impostor score distributions. It reviews biometric verification and identification terminology, simulates 256 Beta-distributed identifiability matrices, and shows that EER can be near zero while Idiff and Dself vary over a wide range, and vice versa. The authors conclude that Idiff and Dself should not be used alone in fingerprinting claims and recommend standard metrics such as rank-k identification rate or EER.","tokens_in":6232,"tokens_out":8126,"duration_ms":80344,"significance":"The mathematical observation is correct and the simulation provides a clean counterexample to the interpretation of Idiff and Dself as recognition-performance measures. The paper contributes a useful terminological and conceptual cleanup by linking neuroscience practice to established biometric standards, and the recommendation to report rank-k or EER as primary metrics is sensible. The availability of simulation code is a strength, and the authors' modal language ('could lead', 'may provide') is appropriately cautious. The main limitation is that the practical prevalence of the misleading cases is not quantified with real data, but this does not affect the logical validity of the core counterexample.","major_comments":[{"comment":"The simulation demonstrates that Idiff and Dself are poorly related to EER, which the authors themselves classify as a verification metric in Section II. However, the conclusion and the practical recommendations are framed around 'fingerprinting' and 'identifiability', i.e., identification, for which the accepted metric is rank-k or the CMC. The current evidence therefore does not directly support the identification-specific claim. Please provide a parallel simulation with rank-1 identification accuracy, or explicitly limit the conclusion to verification contexts and explain why the identification conclusion follows from the same construction.","section":"Section IV, Figures 1-3"},{"comment":"The simulation is not fully described in the text. The number of simulated subjects/identities, the dimension of the identifiability matrices, the procedure for generating diagonal and off-diagonal entries (independence, symmetry, number of Monte Carlo replications), and the method used to compute EER are not specified. These details are necessary for reproducibility; the linked code is welcome, but the text should be self-contained.","section":"Section IV"}],"minor_comments":[{"comment":"The statement 'Including an analysis based on real data seems superfluous in this context' is too dismissive. A real-data demonstration would strengthen the practical relevance of the recommendation, even though the logical counterexample does not require it.","section":"Section V"},{"comment":"The formulas lack parentheses: equation (1) should read (Iself - Iothers) × 100, and equation (2) should read (Corr_ii - μ_ij)/σ_ij.","section":"Equations (1) and (2)"},{"comment":"The Pearson correlation coefficients are reported as '-.707' and '-.536'; the conventional notation is '-0.707' and '-0.536'.","section":"Section IV"},{"comment":"The text says 'Alpha and beta parameters were derived from different mean and standard deviation values to define a total of 256 different scenarios' but does not specify how the mean and standard deviation values were chosen; a brief description or a supplementary table would clarify the design.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the central argument is correct. The missing rank-k simulation and the under-specified simulation details are the main reasons for major revision. No citation or novelty concerns beyond the peripheral self-citation [7], which is not load-bearing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a short methodological commentary on connectome fingerprinting, and its central point is correct: Idiff and Dself are mean-separation statistics, and mean separation alone doesn't determine how much genuine and impostor score distributions overlap. The simulation with 256 Beta-distributed identifiability matrices backs the claim cleanly, showing large spreads of Idiff and Dself at EER values near zero. The code is public, which is good.\n\nWhat's actually new is limited but real: the paper translates a well-known biometric principle—performance depends on overlap, not just mean difference—into a concrete counterexample for two widely used connectome metrics. It also draws a useful terminology boundary between fingerprinting proper (identification/verification with error rates) and exploratory differentiability measures. That part is well argued and cites the relevant biometrics literature.\n\nSoft spots: the authors explicitly decline to run real data ('Including an analysis based on real data seems superfluous in this context'). For the logical point, that's fine—this is an existence/counterexample argument, and the simulated scenarios are legitimate. For the practical conclusion, it leaves open how often real connectome studies are actually misled. If real test-retest data happen to produce distributions where Idiff is strongly ordered with EER, the practical harm would be smaller. The paper's modal language ('could lead', 'may provide') is appropriately restrained, and I don't think the missing real-data check undermines the main claim. It just means the prevalence of the problem is unquantified.\n\nOne small thing: the scatterplots show Pearson correlations of -0.707 and -0.536, which are not as high as the text's 'high correlation' might suggest, and the heteroscedasticity is the real story. The authors do point that out, so it's fine.\n\nCitation pattern: the one self-citation (Demuru & Fraschini) is peripheral. No circular reasoning. The target papers (Idiff from Amico & Goñi, Dself from da Silva Castanheira et al.) are directly addressed.\n\nWho is this for? Researchers who use Idiff or Dself as performance summaries, and reviewers of connectome fingerprinting papers. It doesn't deliver new biology or a new technique, but it does a legitimate corrective job.\n\nRecommendation: worth sending to peer review. A referee with biometrics background can check the simulation and the framing, and may push for a real-data illustration, but the paper is honest and the core argument holds.","headline":"A sound, well-scoped methodological caution: Idiff and Dself are mean-separation statistics that cannot summarize recognition performance, and the Beta simulation makes the counterexample point honestly; absence of real data limits prevalence estimates, not the core argument.","tokens_in":6726,"tokens_out":1673,"would_cite":true,"duration_ms":15341,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mean-difference brain identifiability scores can hide recognition failures","keywords":["brain fingerprinting","identification","verification","connectome","differential identifiability","equal error rate","test-retest reliability","functional connectivity"],"falsifier":"Compute $I_{\\mathrm{diff}}$, $D_{\\mathrm{self}}$, and EER from the identifiability matrix of a large real test-retest dataset with hundreds of participants and two sessions, then count how often EER near zero coincides with low $I_{\\mathrm{diff}}$, or high EER with high $I_{\\mathrm{diff}}$; the paper's claim predicts such cases, and their absence in real data would weaken the practical conclusion.","tokens_in":1904,"feed_emoji":"🧠","tokens_out":4520,"duration_ms":126950,"temperature":0.7,"pith_summary":"The paper argues that the term 'brain fingerprinting' in neuroscience has drifted from its biometric meaning, and that two popular identifiability measures, $I_{\\mathrm{diff}}$ and $D_{\\mathrm{self}}$, are not trustworthy summaries of whether a brain measure can recognize individuals. In 256 simulated scenarios, $I_{\\mathrm{diff}}$ and $D_{\\mathrm{self}}$ correlate only moderately with the biometric Equal Error Rate (EER), and at $\\mathrm{EER}=0$, both measures still range from low to high values. The reason is that both metrics compare average within-subject similarity with average between-subject similarity, ignoring the overlap and shape of the two score distributions that determine recognition errors. The authors conclude that fingerprinting claims should be evaluated with error-based biometric metrics such as EER or rank-k identification, and that 'fingerprinting' should be reserved for evidence of unique, stable individual identification rather than mean differences in similarity.","feed_headline":"Brain fingerprint scores can hide identification failures","feed_subtitle":"Idiff and Dself ignore score overlap, so they can look perfect while recognition errors stay high.","key_machinery":"The carrying object is the identifiability matrix: a square matrix of correlations between every participant's connectivity profile in one session and every participant's profile in a second session. Its main diagonal holds genuine scores (same person), its off-diagonal holds impostor scores (different people). $I_{\\mathrm{diff}}$ is the percentage difference between the average diagonal and average off-diagonal; $D_{\\mathrm{self}}$ is the same idea per subject, expressed as a z-score of a participant's self-correlation relative to their correlations with all others. The paper compares these summaries against the Equal Error Rate (EER), the point where false-accept and false-reject rates are equal, using 256 Beta-distributed scenarios. The mechanism carrying the argument is the contrast between mean separation (what $I_{\\mathrm{diff}}$ and $D_{\\mathrm{self}}$ report) and distribution overlap (what determines EER): two distributions can have the same mean gap yet very different overlap depending on dispersion, skewness, and tails.","core_discovery":"On the paper's own terms, the discovery is that $I_{\\mathrm{diff}}$ and $D_{\\mathrm{self}}$ measure the wrong target. An identifiability matrix contains correlations between every participant's profile at two sessions; diagonal entries are same-person ('genuine') scores and off-diagonal entries are cross-person ('impostor') scores. $I_{\\mathrm{diff}}$ is the percentage difference between the average diagonal and the average off-diagonal, and $D_{\\mathrm{self}}$ is a subject-level z-score of self-correlation relative to correlations with others. In 256 Beta-distributed simulation scenarios, EER correlated with $I_{\\mathrm{diff}}$ at $r=-0.707$ and with average $D_{\\mathrm{self}}$ at $r=-0.536$, yet at $\\mathrm{EER}=0$ both measures varied widely, and similarly at poor EER values. The reason is that mean separation does not capture distribution overlap; because overlap, dispersion, skewness, and threshold determine false-accept and false-reject rates, $I_{\\mathrm{diff}}$ and $D_{\\mathrm{self}}$ can rate a perfectly separable system low and a badly overlapping system high. The paper's positive recommendation is to report EER or rank-k identification rate when the question is identification, and to stop calling mean-difference similarity scores 'fingerprinting.'","pith_inferences":["The paper declines to test real data, so a natural extension is to compute $I_{\\mathrm{diff}}$, $D_{\\mathrm{self}}$, and EER on existing test-retest fMRI/EEG/MEG datasets and measure how often their rankings disagree in practice; the paper's own logic predicts nontrivial disagreement.","If the critique holds, earlier studies that used $I_{\\mathrm{diff}}$ or $D_{\\mathrm{self}}$ as their primary evidence for 'brain fingerprints' would need re-analysis with overlap-sensitive metrics, and some may not support individual identification.","The same overlap argument applies beyond connectomes: any discipline that summarizes two score distributions by the gap between their means—genetic fingerprinting, chemical fingerprinting, device recognition—should prefer error-rate or AUC summaries when the question is recognition.","A testable extension would be constructing a new identifiability index that includes distribution overlap, such as the area under the ROC curve of the identifiability matrix, and seeing whether it reconciles neuroscience practice with biometric standards."],"forward_implications":["Any connectome study reporting only $I_{\\mathrm{diff}}$ or $D_{\\mathrm{self}}$ should not be read as evidence that individuals can be identified; EER, ROC curves, or rank-k rates are needed to support that claim.","Group comparisons of $I_{\\mathrm{diff}}$—for example between patients and controls—are about average similarity differences, not about fingerprinting, and should not be described as identification.","A reported high $I_{\\mathrm{diff}}$ can coexist with high error rates, so results that rely on it alone may need to be rechecked with error-based metrics.","Studies that do use rank-1 identification or EER already satisfy the biometric standard the paper defends, and their 'fingerprinting' terminology is the more appropriate one.","New brain-based identification claims should report thresholds and the success and failure rates they imply, not just score separations."],"supporting_citations":[{"why":"introduced functional connectome fingerprinting with a rank-1 identification rate, the neuroscience origin of the term and an example of proper biometric usage.","marker":"[5]"},{"why":"introduced differential identifiability ($I_{\\mathrm{diff}}$) from the identifiability matrix, the main target metric the paper criticizes.","marker":"[11]"},{"why":"proposed $D_{\\mathrm{self}}$ as a subject-level differentiability measure, the second target metric the paper tests against EER.","marker":"[12]"},{"why":"supplies the biometric definitions of genuine and impostor distributions, FAR, FRR, and recognition that ground the paper's evaluation framework.","marker":"[6]"},{"why":"defines the verification, identification, and authentication terminology the paper uses to argue that many neuroscience uses of 'fingerprinting' are mismatched.","marker":"[4]"}],"fun_headline_variants":["Identifiability scores ignore score overlap","Fingerprint measures can hide high error rates","Mean separation fails to capture identification","Overlap-blind scores distort fingerprint claims","Idiff and Dself miss real identification failures"],"cache_read_input_tokens":9088,"weakest_assumption_plain":"The paper's warning depends on the assumption that Beta-distributed synthetic score distributions cover the realistic shapes of genuine and impostor similarity distributions in real test-retest brain data; if real data constrain those shapes, the practical frequency of disagreement between $I_{\\mathrm{diff}}$ and EER might be lower than the simulation suggests.","fun_headline_variants_meta":{"raw":{"variants":["Identifiability scores ignore score overlap","Fingerprint measures can hide high error rates","Mean separation fails to capture identification","Overlap-blind scores distort fingerprint claims","Idiff and Dself miss real identification failures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1461,"prompt_tokens":994,"completion_tokens":467,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":403}},"tokens_in":610,"tokens_out":467,"duration_ms":4956,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:13:03.757970+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $I_{\\mathrm{diff}}$, $D_{\\mathrm{self}}$, and EER from the identifiability matrix of a large real test-retest dataset with hundreds of participants and two sessions, then count how often EER near zero coincides with low $I_{\\mathrm{diff}}$, or high EER with high $I_{\\mathrm{diff}}$; the paper's claim predicts such cases, and their absence in real data would weaken the practical conclusion.","supporting_citations":[{"cited_title":"da Silva Castanheira, H","cited_arxiv_id":null,"evidence_quote":"proposed $D_{\\mathrm{self}}$ as a subject-level differentiability measure, the second target metric the paper tests against EER."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"supplies the biometric definitions of genuine and impostor distributions, FAR, FRR, and recognition that ground the paper's evaluation framework."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"defines the verification, identification, and authentication terminology the paper uses to argue that many neuroscience uses of 'fingerprinting' are mismatched."}],"review_version":1}