Pith. sign in

REVIEW 2 major objections 2 minor 166 references

Provable Joint Decontamination for Benchmarking Multiple Large Language Models

T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3

Pith's one-line read A conformal procedure selects shared benchmarks for multiple LLMs while provably controlling the global contamination rate.

desk verdict JECS formalizes joint multi-model decontamination via max-p aggregation and right-tail envelope reconstruction for GCR control, but the envelope may lose conservativeness under realistic cross-model dependence. read the letter →

arxiv 2605.21543 v1 pith:75LO3J3K submitted 2026-05-20 cs.LG

classification cs.LG
keywords benchmarkcontaminationLLMevaluationconformalselectionglobalratemultipletestingBenjamini-Hochberg
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Benchmark contamination inflates LLM performance scores when test examples appear in training data, making cross-model comparisons unreliable. The paper formalizes this as a joint selection problem and introduces Joint Envelope Conformal Selection (JECS). JECS calculates per-model conformal p-values, aggregates by taking the maximum per item, and builds a conservative envelope of the null distribution using right-tail observations above a threshold. It then applies an adaptive Benjamini-Hochberg procedure to the rescaled values to select items with guaranteed control on the overall contamination fraction. Experiments confirm higher power than baselines while keeping the global rate at the target level.

What carries the argument

The reconstruction of a conservative envelope for the distribution of maximum p-values, which rescales the statistics to enable joint control across models.

What would settle it

A simulation or real audit in which the fraction of actually contaminated items among those selected by JECS exceeds the nominal global contamination rate target.

Watch

Extended reading notes

Core claim

JECS computes per-model conformal p-values, aggregates them by the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure to select a benchmark with provable global contamination rate control.

Load-bearing premise

The conservative envelope built from right-tail observations above a data-driven threshold continues to upper-bound the true max-p null distribution when the individual conformal p-values meet their stated assumptions.

Editorial extensions

If this is right

  • Provides a single decontaminated benchmark that supports fair performance comparisons across all audited models.
  • Maintains the target global contamination rate control even when separate per-model detections would produce inconsistent selections.
  • Achieves higher statistical power for identifying clean items compared to applying the maximum p-value directly without the envelope adjustment.
  • Applies to various LLMs and benchmarks while consistently respecting the contamination control target.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Practitioners could adopt this for standardized evaluation suites shared among multiple model developers to reduce hidden contamination effects.
  • The envelope approach might generalize to other joint multiple-testing problems where the maximum statistic is of interest.
  • Testing the method on benchmarks with known partial contamination could reveal how sensitive the data-driven threshold is in practice.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper formalizes multi-model benchmark decontamination as a joint selection problem and proposes Joint Envelope Conformal Selection (JECS). JECS computes per-model conformal p-values, aggregates them via the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure to the rescaled values to select a benchmark with provable global contamination rate (GCR) control under stated assumptions. Experiments across models and benchmarks are reported to show higher power than the max-p baseline while maintaining target GCR.

Significance. If the GCR control holds under the procedure's assumptions, the result would be significant for LLM evaluation: it would provide the first joint conformal method that enables fair cross-model comparisons by producing a single decontaminated benchmark with theoretical guarantees, rather than model-specific selections. The use of conformal p-values and adaptive BH is a natural extension of existing single-model decontamination work.

major comments (2)
  1. The central claim of provable GCR control rests on the conservative envelope reconstruction step, but the manuscript provides no explicit derivation, no statement of the full set of assumptions (e.g., on the dependence structure among per-model p-values), and no quantitative experimental results or tables in the provided description. This makes the guarantee impossible to verify from the text and is load-bearing for the main contribution.
  2. The envelope reconstruction from right-tail observations above a data-driven threshold (method description) may fail to stochastically dominate the true null distribution of the maximum when per-model conformal p-values for the same item are positively dependent, which is realistic for shared benchmark items. The data-driven threshold correlates with the observed maxima, and no explicit modeling or robustness argument for this dependence is given; a counterexample or additional proof under dependence would be required to support the GCR claim.
minor comments (2)
  1. The abstract claims 'extensive experiments' with higher power and GCR control but reports no numerical values, tables, or specific benchmarks/models; adding these would improve clarity.
  2. Notation for the envelope and rescaled p-values should be defined more precisely with equation numbers to allow direct reference in the proof.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the thoughtful and detailed report. The comments correctly identify that the theoretical guarantees are central to the contribution, and we will revise the manuscript to make the derivation, assumptions, and empirical verification fully explicit and self-contained.

read point-by-point responses
  1. Referee: The central claim of provable GCR control rests on the conservative envelope reconstruction step, but the manuscript provides no explicit derivation, no statement of the full set of assumptions (e.g., on the dependence structure among per-model p-values), and no quantitative experimental results or tables in the provided description. This makes the guarantee impossible to verify from the text and is load-bearing for the main contribution.

    Authors: We agree that the current presentation would benefit from greater explicitness. In the revised manuscript we will add a dedicated subsection (and appendix) containing the full derivation of the conservative envelope, the precise statement of all assumptions (including marginal validity of the per-model conformal p-values and the conditions under which the right-tail envelope stochastically dominates the null distribution of the maximum), and a self-contained proof of GCR control. We will also expand the experimental section with additional tables that report empirical GCR on both real benchmarks and controlled synthetic dependence scenarios. revision: yes

  2. Referee: The envelope reconstruction from right-tail observations above a data-driven threshold (method description) may fail to stochastically dominate the true null distribution of the maximum when per-model conformal p-values for the same item are positively dependent, which is realistic for shared benchmark items. The data-driven threshold correlates with the observed maxima, and no explicit modeling or robustness argument for this dependence is given; a counterexample or additional proof under dependence would be required to support the GCR claim.

    Authors: The concern is well-taken. Positive dependence among p-values for the same item is plausible and the data-driven threshold introduces correlation with the observed maxima. In the revision we will (i) explicitly state the working assumption of conditional independence across models given the item (or, alternatively, provide a worst-case bound), (ii) supply a short proof sketch showing that the envelope remains conservative under this assumption, and (iii) include a small simulation study that injects controlled positive dependence and verifies that the realized GCR stays below the nominal level. If the dependence is judged too strong for the guarantee to hold, we will clearly delineate the limitation in the revised text. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

JECS derivation relies on standard conformal and BH theory with independent envelope construction

full rationale

The paper proposes JECS by computing per-model conformal p-values (standard), aggregating via per-item max, reconstructing a conservative envelope of the max-p null from right-tail observations above a data-driven threshold, and feeding rescaled values into adaptive BH for GCR control. This chain builds on established conformal inference and multiple-testing results without reducing the claimed guarantee to a tautology, self-definition, or fitted input renamed as prediction. The envelope step is presented as a novel but externally motivated construction under stated assumptions rather than presupposing the target GCR; no load-bearing premise collapses to self-citation or ansatz smuggling. The procedure remains self-contained against external benchmarks for its core components.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The procedure rests on standard conformal prediction assumptions plus the validity of the right-tail envelope reconstruction; no free parameters or invented entities are mentioned.

assumptions (2)
  • domain assumption Per-model conformal p-values are valid under the null of no contamination for each model.
    Invoked when computing per-model p-values before aggregation.
  • domain assumption The max-p null distribution can be conservatively reconstructed from right-tail observations above a data-driven threshold.
    Central to the envelope construction that delivers GCR control.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Provable Joint Decontamination for Benchmarking Multiple Large Language Models." pith.science (2026). https://pith.science/paper/75LO3J3K

@misc{pith2026260521543,
  author       = {Pith},
  title        = {Pith review of: Provable Joint Decontamination for Benchmarking Multiple Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/75LO3J3K}},
  note         = {Machine review of arXiv:2605.21543}
}
read the original abstract

Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performance can be inflated and cross-model comparisons become unreliable. A broad line of training-data detection work designs scores to quantify how strongly a model memorizes a given data point, but these score-based methods lack theoretical guarantees. Recent conformal approaches provide provable false-identification control for a single model; however, applying them separately to each model can produce model-specific benchmarks, undermining fair comparison across models. In this work, we formalize multi-model benchmark decontamination as a joint selection problem and propose Joint Envelope Conformal Selection (JECS), a conformal procedure that enables global contamination rate (GCR) control under stated assumptions. Specifically, JECS computes per-model conformal p-values, aggregates them by the per-item maximum, and reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold. By applying the adaptive Benjamini-Hochberg (BH) procedure to the envelope-rescaled values, we select a benchmark with provable GCR control. Extensive experiments across various models and benchmarks demonstrate that JECS achieves higher power than the max-p baseline while consistently maintaining the target GCR control.

Figures

Figures reproduced from arXiv: 2605.21543 by the authors.

Figure 1
Figure 1. Null density of the per-model conformal p versus p ∗ i = maxk p k i at K = 8. The super-uniformity tax. JMCS is valid in finite samples but substantially conservative. The BH cutoff αr/n is calibrated to the uniform scale, whereas under H0,i the maximum statistic p ∗ i = maxk p k i is strictly super-uniform: in the stylized homogeneous null with p 1 i , . . . , pK i iid∼ Unif(0, 1), the null CDF is F0(t) = t K and t… view at source ↗
Figure 2
Figure 2. Fitted envelope schematic. The fitted Fbfit is a conservative estimate of the true F0. The left branch is linear with slope Fbfit(λ)/λ, whereas the right branch uses the empirical right￾tail CDF rescaled to match the anchor [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Joint GCR control with JECS at K = 16. Each panel corresponds to a dataset–model family pair and reports realized GCP curves on the left axis and Power bars on the right axis for four detection scores under JECS; the dashed diagonal is the target GCR. Detection scores and baseline. We employ the per-model score T(x; θk) with four standard detectors: Perplexity [Carlini et al., 2021], Min-K% [Shi et al., 2024], Min-K… view at source ↗
Figures from the paper (5 more)
Figure 5
Figure 5. Figure 5: GCP curve (left axis) and Power bars (right axis) under varying training fraction ρ. curve is strictly monotone in K, but that the procedure remains controlled as the number of jointly audited models grows; individual power trends can be non-monotone because, by Equati…
Figure 6
Figure 6. Figure 6: Sensitivity to the number of audited models [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Realized GCP and Power of JECS on five MIMIR subsets at [PITH_FULL_IMAGE:figures/full_fig_p021_7.png]
Figure 8
Figure 8. Figure 8: JECS with Min-K%++ at K = 16: cross-experiment dispersion. Curves report mean realized GCP (left axis, orange) and Power (right axis, blue) on WikiMIA (top row) and ArXivTection (bottom row), for GPT-NeoX-20B, LLaMA-7B, and Pythia-6.9B. Shaded bands are ±1 standard dev…
Figure 9
Figure 9. Figure 9: reports realized GCP and Power on this mixed-family pool for the four detection scores at K = 16, on WikiMIA and ArXivTection. The realized GCP stays at or below the target diagonal throughout; Min-K%++ delivers the highest Power on ArXivTection while the four scores c…

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • IndisputableMonolith/Cost/FunctionalEquation.lean washburn_uniqueness_aczel unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    JECS computes per-model conformal p-values, aggregates them by the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Reference graph

Works this paper leans on

166 extracted references · 166 canonical work pages

  1. [1]

    2025 , month =

    Complaint for Copyright Infringement , howpublished =. 2025 , month =

  2. [2]

    IEEE Symposium on Security and Privacy , pages=

    Membership inference attacks against machine learning models , author=. IEEE Symposium on Security and Privacy , pages=. 2017 , organization=

  3. [3]

    International Conference on Learning Representations , year=

    How much of my dataset did you use? Quantitative Data Usage Inference in Machine Learning , author=. International Conference on Learning Representations , year=

  4. [4]

    International Conference on Learning Representations , year=

    Quantifying Memorization Across Neural Language Models , author=. International Conference on Learning Representations , year=

  5. [5]

    2016 , month =

    Regulation (EU) 2016/679 (General Data Protection Regulation) , howpublished =. 2016 , month =

  6. [6]

    2018 , number =

    The California Consumer Privacy Act of 2018 (CCPA) , institution =. 2018 , number =

  7. [7]

    IEEE Conference on Secure and Trustworthy Machine Learning , pages=

    Position: Membership Inference Attacks Cannot Prove That a Model was Trained on Your Data , author=. IEEE Conference on Secure and Trustworthy Machine Learning , pages=. 2025 , organization=

  8. [8]

    arXiv preprint arXiv:2309.10677 , year=

    Estimating contamination via perplexity: Quantifying memorisation in language model evaluation , author=. arXiv preprint arXiv:2309.10677 , year=

Show all 166 references
  1. [9]

    International Conference on Learning Representations , year=

    Detecting Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=

  2. [10]

    IEEE Symposium on Security and Privacy (SP) , pages=

    Membership inference attacks from first principles , author=. IEEE Symposium on Security and Privacy (SP) , pages=. 2022 , organization=

  3. [11]

    2018 , organization=

    Privacy risk in machine learning: Analyzing the connection to overfitting , author=. 2018 , organization=

  4. [12]

    Ahmed Salem and Yang Zhang and Mathias Humbert and Pascal Berrang and Mario Fritz and Michael Backes , title =

  5. [13]

    ACM SIGSAC Conference on Computer and Communications Security , pages=

    Enhanced membership inference attacks against machine learning models , author=. ACM SIGSAC Conference on Computer and Communications Security , pages=

  6. [14]

    Advances in Neural Information Processing Systems , volume=

    Scalable membership inference attacks via quantile regression , author=. Advances in Neural Information Processing Systems , volume=

  7. [15]

    The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

    Membership Inference Attacks against Large Vision-Language Models , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=

  8. [16]

    International Conference on Machine Learning , year=

    Low-Cost High-Power Membership Inference Attacks , author=. International Conference on Machine Learning , year=

  9. [17]

    Advances in Neural Information Processing Systems , year=

    Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration , author=. Advances in Neural Information Processing Systems , year=

  10. [18]

    USENIX Security Symposium , pages=

    Extracting training data from large language models , author=. USENIX Security Symposium , pages=

  11. [19]

    arXiv preprint arXiv:2503.07482 , year=

    Efficient Membership Inference Attacks by Bayesian Neural Network , author=. arXiv preprint arXiv:2503.07482 , year=

  12. [20]

    International Conference on Learning Representations , year=

    Fine-tuning can Help Detect Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=

  13. [21]

    Journal of Machine Learning Research , volume=

    Selection by prediction with conformal p-values , author=. Journal of Machine Learning Research , volume=

  14. [22]

    Journal of the Royal Statistical Society , volume=

    Controlling the false discovery rate: a practical and powerful approach to multiple testing , author=. Journal of the Royal Statistical Society , volume=. 1995 , publisher=

  15. [23]

    Annals of Statistics , pages=

    The control of the false discovery rate in multiple testing under dependency , author=. Annals of Statistics , pages=. 2001 , publisher=

  16. [24]

    International Conference on Learning Representations , year=

    Proving test set contamination in black-box language models , author=. International Conference on Learning Representations , year=

  17. [25]

    International Conference on Learning Representations , year=

    A Statistical Approach for Controlled Training Data Detection , author=. International Conference on Learning Representations , year=

  18. [26]

    Deyao Zhu and Jun Chen and Xiaoqian Shen and Xiang Li and Mohamed Elhoseiny , booktitle=. Mini. 2024 , url=

  19. [27]

    Advances in Neural Information Processing Systems , volume=

    Visual instruction tuning , author=. Advances in Neural Information Processing Systems , volume=

  20. [29]

    International Conference on Machine Learning , pages=

    Pythia: A suite for analyzing large language models across training and scaling , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  21. [30]

    arXiv preprint arXiv:2205.01068 , year=

    Opt: Open pre-trained transformer language models , author=. arXiv preprint arXiv:2205.01068 , year=

  22. [32]

    IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Deep residual learning for image recognition , author=. IEEE Conference on Computer Vision and Pattern Recognition , pages=

  23. [33]

    IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Densely connected convolutional networks , author=. IEEE Conference on Computer Vision and Pattern Recognition , pages=

  24. [34]

    International Conference on Machine Learning , year=

    Andr. International Conference on Machine Learning , year=

  25. [35]

    2019 , url=

    Language Models are Unsupervised Multitask Learners , author=. 2019 , url=

  26. [36]

    arXiv preprint arXiv:2101.00027 , year=

    The Pile: An 800GB Dataset of Diverse Text for Language Modeling , author=. arXiv preprint arXiv:2101.00027 , year=

  27. [37]

    and Lapata, Mirella

    Narayan, Shashi and Cohen, Shay B. and Lapata, Mirella. Don ' t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/D18-1206

  28. [38]

    Yucheng Li and Frank Guerin and Chenghua Lin , title =

  29. [39]

    2009 , journal=

    Learning multiple layers of features from tiny images , author=. 2009 , journal=

  30. [40]

    USENIX Security Symposium , pages=

    Systematic evaluation of privacy risks of machine learning models , author=. USENIX Security Symposium , pages=

  31. [41]

    Berkeley Symposium on Mathematical Statistics and Probability , volume=

    On measures of entropy and information , author=. Berkeley Symposium on Mathematical Statistics and Probability , volume=. 1961 , organization=

  32. [42]

    Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models

    Dong, Yihong and Jiang, Xue and Liu, Huanyu and Jin, Zhi and Gu, Bin and Yang, Mengfei and Li, Ge. Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024...

  33. [45]

    doi:10.5281/zenodo.12608602 , url =

    Gao, Leo and Tow, Jonathan and Abbasi, Baber and Biderman, Stella and Black, Sid and DiPofi, Anthony and Foster, Charles and Golding, Laurence and Hsu, Jeffrey and Le Noac'h, Alain and Li, Haonan and McDonell, Kyle and Muennighoff, Niklas and Ociepa, Chris and Phang, Jason and...

  34. [46]

    NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

    Sainz, Oscar and Campos, Jon and Garc \'i a-Ferrero, Iker and Etxaniz, Julen and de Lacalle, Oier Lopez and Agirre, Eneko. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. Findings of the Association for Computational Linguistics: EM...

  35. [47]

    Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLM s

    Balloccu, Simone and Schmidtov \'a , Patr \'i cia and Lango, Mateusz and Dusek, Ondrej. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLM s. Proceedings of the 18th Conference of the European Chapter of the Association for Computational L...

  36. [48]

    Proceedings of the IEEE Symposium on Security and Privacy (SP) , pages=

    Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning , author=. Proceedings of the IEEE Symposium on Security and Privacy (SP) , pages=. 2019 , organization=

  37. [49]

    International Conference on Machine Learning , pages=

    White-box vs black-box: Bayes optimal strategies for membership inference , author=. International Conference on Machine Learning , pages=. 2019 , organization=

  38. [50]

    Advances in Neural Information Processing Systems , year =

    Gaussian Membership Inference Privacy , author =. Advances in Neural Information Processing Systems , year =

  39. [51]

    IEEE Transactions on Computational Social Systems , volume=

    Socinf: Membership inference attacks on social media health data with machine learning , author=. IEEE Transactions on Computational Social Systems , volume=. 2019 , publisher=

  40. [52]

    Lauren Watson and Chuan Guo and Graham Cormode and Alexandre Sablayrolles , title =

  41. [54]

    Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =

    Inbal Magar and Roy Schwartz , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =

  42. [55]

    International Conference on Learning Representations , year=

    Min-K\ author=. International Conference on Learning Representations , year=

  43. [56]

    2024 , booktitle=

    Do Membership Inference Attacks Work on Large Language Models? , author=. 2024 , booktitle=

  44. [58]

    International Conference on Learning Representations , year=

    Infilling Score: A Pretraining Data Detection Algorithm for Large Language Models , author=. International Conference on Learning Representations , year=

  45. [60]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=

    A direct approach to false discovery rates , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2002 , publisher=

  46. [61]

    Vladimir Vovk and Ilia Nouretdinov and Alex Gammerman , title =

  47. [62]

    2005 , publisher=

    Algorithmic learning in a random world , author=. 2005 , publisher=

  48. [63]

    The Annals of Statistics , volume=

    Testing for outliers with conformal p-values , author=. The Annals of Statistics , volume=. 2023 , publisher=

  49. [64]

    Computational Statistics & Data Analysis , volume=

    Beta kernel estimators for density functions , author=. Computational Statistics & Data Analysis , volume=. 1999 , publisher=

  50. [65]

    Journal of Nonparametric Statistics , volume=

    Bias reductions for beta kernel estimation , author=. Journal of Nonparametric Statistics , volume=. 2016 , publisher=

  51. [66]

    Journal of the American Statistical Association , volume=

    Improvement of kernel type density estimators , author=. Journal of the American Statistical Association , volume=. 1977 , publisher=

  52. [67]

    Pierre Neuvial , title =. J. Mach. Learn. Res. , volume =

  53. [68]

    Biometrika , volume=

    Adaptive linear step-up procedures that control the false discovery rate , author=. Biometrika , volume=. 2006 , publisher=

  54. [69]

    Journal of Educational and Behavioral Statistics , volume=

    On the adaptive control of the false discovery rate in multiple testing with independent statistics , author=. Journal of Educational and Behavioral Statistics , volume=. 2000 , publisher=

  55. [70]

    Proceedings of the Sixteenth International Conference on Machine Learning , pages =

    Vovk, Volodya and Gammerman, Alexander and Saunders, Craig , title =. Proceedings of the Sixteenth International Conference on Machine Learning , pages =. 1999 , isbn =

  56. [71]

    ACM transactions on intelligent systems and technology , volume=

    A survey on evaluation of large language models , author=. ACM transactions on intelligent systems and technology , volume=. 2024 , publisher=

  57. [72]

    Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

    Investigating data contamination in modern benchmarks for large language models , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=

  58. [74]

    Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

    NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=

  59. [76]

    Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs , author=. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  60. [78]

    arXiv preprint arXiv:2211.15533 , year=

    The stack: 3 tb of permissively licensed source code , author=. arXiv preprint arXiv:2211.15533 , year=

  61. [79]

    arXiv preprint arXiv:2305.06161 , year=

    Starcoder: may the source be with you! , author=. arXiv preprint arXiv:2305.06161 , year=

  62. [80]

    Advances in Neural Information Processing Systems , volume=

    Redpajama: an open dataset for training large language models , author=. Advances in Neural Information Processing Systems , volume=

  63. [82]

    Advances in Neural Information Processing Systems , volume=

    The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only , author=. Advances in Neural Information Processing Systems , volume=

  64. [83]

    2024 , month = jul, day =

  65. [84]

    Forty-third International Conference on Machine Learning , year=

    Provable Training Data Identification for Large Language Models , author=. Forty-third International Conference on Machine Learning , year=

  66. [85]

    The Fourteenth International Conference on Learning Representations , year=

    Multi-Condition Conformal Selection , author=. The Fourteenth International Conference on Learning Representations , year=

  67. [86]

    Forty-second International Conference on Machine Learning , year=

    Multivariate Conformal Selection , author=. Forty-second International Conference on Machine Learning , year=

  68. [87]

    2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=

    Sok: Membership inference attacks on llms are rushing nowhere (and how to fix it) , author=. 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=. 2025 , organization=

  69. [88]

    Advances in Neural Information Processing Systems , volume=

    LLM Dataset Inference: Did you train on my dataset? , author=. Advances in Neural Information Processing Systems , volume=

  70. [89]

    Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=

    Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2004 , publisher=

  71. [90]

    Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

    Detecting Non-Membership in LLM Training Data via Rank Correlations , author=. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

  72. [91]

    Forty-second International Conference on Machine Learning , year=

    How Contaminated Is Your Benchmark? Measuring Dataset Leakage in Large Language Models with Kernel Divergence , author=. Forty-second International Conference on Machine Learning , year=

  73. [92]

    The Fourteenth International Conference on Learning Representations , year=

    BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models , author=. The Fourteenth International Conference on Learning Representations , year=

  74. [94]

    The Twelfth International Conference on Learning Representations , year=

    Proving Test Set Contamination in Black-Box Language Models , author=. The Twelfth International Conference on Learning Representations , year=

  75. [97]

    2026 , url=

    Yihao LIU and Xinqi LYU and Dong Wang and Yanjie Li and Bin Xiao , booktitle=. 2026 , url=

  76. [98]

    Yuke Hu and Zheng Li and Zhihao Liu and Yang Zhang and Zhan Qin and Kui Ren and Chun Chen , title =

  77. [99]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Practical membership inference attacks against large-scale multi-modal models: A pilot study , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  78. [101]

    2026 , eprint=

    Membership Inference on LLMs in the Wild , author=. 2026 , eprint=

  79. [102]

    Detecting Data Contamination in

    Micha. Detecting Data Contamination in. The Fourteenth International Conference on Learning Representations , year=

  80. [104]

    The Twelfth International Conference on Learning Representations , year=

    DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks , author=. The Twelfth International Conference on Learning Representations , year=

  81. [106]

    Koala: An Index for Quantifying Overlaps with Pre-training Corpora , booktitle =

    Thuy. Koala: An Index for Quantifying Overlaps with Pre-training Corpora , booktitle =

  82. [107]

    GitHub repository , howpublished =

    Yanyang Li , title =. GitHub repository , howpublished =. 2024 , publisher =

  83. [108]

    LiveBench: A Challenging, Contamination-Limited

    Colin White and Samuel Dooley and Manley Roberts and Arka Pal and Benjamin Feuer and Siddhartha Jain and Ravid Shwartz-Ziv and Neel Jain and Khalid Saifullah and Sreemanti Dey and Shubh-Agrawal and Sandeep Singh Sandha and Siddartha Venkat Naidu and Chinmay Hegde and Yann LeCu...

  84. [109]

    2024 , eprint=

    C ^2 LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation , author=. 2024 , eprint=

  85. [110]

    Sebastian Bordt and Suraj Srinivas and Valentyn Boreiko and Ulrike von Luxburg , title =

  86. [111]

    2008 , publisher=

    Inductive conformal prediction: Theory and application to neural networks , author=. 2008 , publisher=

  87. [112]

    Biometrika , pages=

    Model-free selective inference under covariate shift via weighted conformal p-values , author=. Biometrika , pages=. 2025 , publisher=

  88. [113]

    Advances in Neural Information Processing Systems , volume=

    Conformal alignment: Knowing when to trust foundation models with guarantees , author=. Advances in Neural Information Processing Systems , volume=

  89. [114]

    Journal of Chemical Information and Modeling , volume=

    Conformal selection for efficient and accurate compound screening in drug discovery , author=. Journal of Chemical Information and Modeling , volume=. 2025 , publisher=

  90. [115]

    Time Travel in

    Shahriar Golchin and Mihai Surdeanu , booktitle=. Time Travel in. 2024 , url=

  91. [116]

    International Conference on Learning Representations , year=

    Fine-tuning Can Help Detect Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=

  92. [117]

    Gpt-4 technical report

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  93. [118]

    Conformal selection for efficient and accurate compound screening in drug discovery

    Tian Bai, Peng Tang, Yuting Xu, Vladimir Svetnik, Bingjia Yang, Abbas Khalili, Xiang Yu, and Archer Y Yang. Conformal selection for efficient and accurate compound screening in drug discovery. Journal of Chemical Information and Modeling, 65 0 (24): 0 13070--13085, 2025 a

  94. [119]

    Tian Bai, Yue Zhao, Xiang Yu, and Archer Y. Yang. Multivariate conformal selection. In Forty-second International Conference on Machine Learning, 2025 b . URL https://openreview.net/forum?id=g2tr7nA4pS

  95. [120]

    Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

    Simone Balloccu, Patr \' cia Schmidtov \'a , Mateusz Lango, and Ond r ej Du s ek. Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Lingu...

  96. [121]

    Testing for outliers with conformal p-values

    Stephen Bates, Emmanuel Cand \`e s, Lihua Lei, Yaniv Romano, and Matteo Sesia. Testing for outliers with conformal p-values. The Annals of Statistics, 51 0 (1): 0 149--178, 2023

  97. [122]

    Controlling the false discovery rate: a practical and powerful approach to multiple testing

    Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, 57 0 (1): 0 289--300, 1995

  98. [123]

    The control of the false discovery rate in multiple testing under dependency

    Yoav Benjamini and Daniel Yekutieli. The control of the false discovery rate in multiple testing under dependency. Annals of Statistics, pages 1165--1188, 2001

  99. [124]

    Pythia: A suite for analyzing large language models across training and scaling

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In Intern...

  100. [125]

    GPT - N eo X -20 B : An open-source autoregressive language model

    Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. GPT - N eo X -20 ...

  101. [126]

    How much can we forget about data contamination? In ICML , Proceedings of Machine Learning Research

    Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko, and Ulrike von Luxburg. How much can we forget about data contamination? In ICML , Proceedings of Machine Learning Research. PMLR / OpenReview.net, 2025

  102. [127]

    Extracting training data from large language models

    Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In USENIX Security Symposium, pages 2633--2650, 2021

  103. [128]

    A survey on evaluation of large language models

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15 0 (3): 0 1--45, 2024

  104. [129]

    A survey on data contamination for large language models

    Yuxing Cheng, Yi Chang, and Yuan Wu. A survey on data contamination for large language models. arXiv preprint arXiv:2502.14425, 2025

  105. [130]

    How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence

    Hyeong Kyu Choi, Maxim Khanov, Hongxin Wei, and Yixuan Li. How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=wVDR2qmE28

  106. [131]

    Investigating data contamination in modern benchmarks for large language models

    Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. Investigating data contamination in modern benchmarks for large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human...

  107. [132]

    Do membership inference attacks work on large language models? In Conference on Language Modeling, 2024

    Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. Do membership inference attacks work on large language models? In Conference on Language Modeling, 2024

  108. [133]

    Oliveira, and Lei Li

    Andr \'e Vicente Duarte, Xuandong Zhao, Arlindo L. Oliveira, and Lei Li. DE - COP : Detecting copyrighted content in language models training data. In International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=LO4xhXmFal

  109. [134]

    Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)

    European Parliament and Council of the European Union . Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) . Official Journal of the European Union, OJ L, 2024/1689, 12 J...

  110. [135]

    Time travel in LLM s: Tracing data contamination in large language models

    Shahriar Golchin and Mihai Surdeanu. Time travel in LLM s: Tracing data contamination in large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=2Rwq6c3tvr

  111. [136]

    Conformal alignment: Knowing when to trust foundation models with guarantees

    Yu Gui, Ying Jin, and Zhimei Ren. Conformal alignment: Knowing when to trust foundation models with guarantees. Advances in Neural Information Processing Systems, 37: 0 73884--73919, 2024

  112. [137]

    Multi-condition conformal selection

    Qingyang Hao, Wenbo Liao, Bingyi Jing, and Hongxin Wei. Multi-condition conformal selection. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=giL8Q1V26J

  113. [138]

    Membership inference attacks against vision-language models

    Yuke Hu, Zheng Li, Zhihao Liu, Yang Zhang, Zhan Qin, Kui Ren, and Chun Chen. Membership inference attacks against vision-language models. In USENIX Security Symposium , pages 1589--1608. USENIX Association, 2025 a

  114. [139]

    A statistical approach for controlled training data detection

    Zirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen, and Dacheng Tao. A statistical approach for controlled training data detection. In International Conference on Learning Representations, 2025 b . URL https://openreview.net/forum?id=XAN8G0rvoB

  115. [140]

    Selection by prediction with conformal p-values

    Ying Jin and Emmanuel J Cand \`e s. Selection by prediction with conformal p-values. Journal of Machine Learning Research, 24 0 (244): 0 1--41, 2023

  116. [141]

    Model-free selective inference under covariate shift via weighted conformal p-values

    Ying Jin and Emmanuel J Cand \`e s. Model-free selective inference under covariate shift via weighted conformal p-values. Biometrika, page asaf066, 2025

  117. [142]

    Sampling-based pseudo-likelihood for membership inference attacks

    Masahiro Kaneko, Youmi Ma, Yuki Wata, and Naoaki Okazaki. Sampling-based pseudo-likelihood for membership inference attacks. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL ...

  118. [143]

    Scaling laws for neural language models

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020

  119. [144]

    Practical membership inference attacks against large-scale multi-modal models: A pilot study

    Myeongseob Ko, Ming Jin, Chenguang Wang, and Ruoxi Jia. Practical membership inference attacks against large-scale multi-modal models: A pilot study. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4871--4881, 2023

  120. [145]

    Task contamination: language models may not be few-shot anymore

    Changmao Li and Jeffrey Flanigan. Task contamination: language models may not be few-shot anymore. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Sy...

  121. [146]

    Awesome data contamination

    Yanyang Li. Awesome data contamination. https://github.com/lyy1994/awesome-data-contamination, 2024

  122. [147]

    Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction

    Yucheng Li, Frank Guerin, and Chenghua Lin. Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction. In AAAI Conference on Artificial Intelligence , pages 18600--18607. AAAI Press, 2024 a

  123. [148]

    An open-source data contamination report for large language models

    Yucheng Li, Yunhao Guo, Frank Guerin, and Chenghua Lin. An open-source data contamination report for large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 528--541, Mia...

  124. [149]

    Membership inference attacks against large vision-language models

    Zhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin, Elias Abad Rocamora, and Volkan Cevher. Membership inference attacks against large vision-language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 c . URL https://openreview.net/fo...

  125. [150]

    LOMIA : Label-only membership inference attacks against pre-trained large vision-language models

    Yihao LIU, Xinqi LYU, Dong Wang, Yanjie Li, and Bin Xiao. LOMIA : Label-only membership inference attacks against pre-trained large vision-language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. URL https://openreview.net/forum?id...

  126. [151]

    Provable training data identification for large language models

    Zhenlong Liu, Hao Zeng, Weiran Huang, and Hongxin Wei. Provable training data identification for large language models. In Forty-third International Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2510.09717. arXiv preprint arXiv:2510.09717

  127. [152]

    Llm dataset inference: Did you train on my dataset? Advances in Neural Information Processing Systems, 37: 0 124069--124092, 2024

    Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. Llm dataset inference: Did you train on my dataset? Advances in Neural Information Processing Systems, 37: 0 124069--124092, 2024

  128. [153]

    Membership inference attacks against language models via neighbourhood comparison

    Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. Membership inference attacks against language models via neighbourhood comparison. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findi...

  129. [154]

    Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

    Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. Proving test set contamination in black-box language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=KS8mIvetg2

  130. [155]

    Inductive conformal prediction: Theory and application to neural networks

    Harris Papadopoulos. Inductive conformal prediction: Theory and application to neural networks. INTECH Open Access Publisher Rijeka, 2008

  131. [156]

    The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only. Advances in Neural Inf...

  132. [157]

    Infilling score: A pretraining data detection algorithm for large language models

    Negin Raoof, Litu Rout, Giannis Daras, Sujay Sanghavi, Constantine Caramanis, Sanjay Shakkottai, and Alex Dimakis. Infilling score: A pretraining data detection algorithm for large language models. In International Conference on Learning Representations, 2025. URL https://open...

  133. [158]

    Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

    Oscar Sainz, Jon Campos, Iker Garc \' a-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages...

  134. [159]

    Detecting pretraining data from large language models

    Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zWqr3MQuNs

  135. [160]

    Membership inference attacks against machine learning models

    Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In IEEE Symposium on Security and Privacy, pages 3--18. IEEE, 2017

  136. [161]

    Systematic evaluation of privacy risks of machine learning models

    Liwei Song and Prateek Mittal. Systematic evaluation of privacy risks of machine learning models. In USENIX Security Symposium, pages 2615--2632, 2021

  137. [162]

    Beyondbench: Contamination-resistant evaluation of reasoning in language models

    Gaurav Srivastava, Aafiya Shamshad Hussain, Zhenyu Bi, Swastik Roy, Priya Pitre, Meng Lu, Morteza Ziyadi, and Xuan Wang. Beyondbench: Contamination-resistant evaluation of reasoning in language models. In The Fourteenth International Conference on Learning Representations, 202...

  138. [163]

    A direct approach to false discovery rates

    John D Storey. A direct approach to false discovery rates. Journal of the Royal Statistical Society Series B: Statistical Methodology, 64 0 (3): 0 479--498, 2002

  139. [164]

    Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach

    John D Storey, Jonathan E Taylor, and David Siegmund. Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach. Journal of the Royal Statistical Society Series B: Statistical Methodology, 66 0 (1): 0 1...

  140. [165]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023

  141. [166]

    Testing exchangeability on-line

    Vladimir Vovk, Ilia Nouretdinov, and Alex Gammerman. Testing exchangeability on-line. In International Conference on Machine Learning , pages 768--775. AAAI Press, 2003

  142. [167]

    Algorithmic learning in a random world

    Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic learning in a random world. Springer, 2005

  143. [168]

    Koala: An index for quantifying overlaps with pre-training corpora

    Thuy - Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi. Koala: An index for quantifying overlaps with pre-training corpora. In EMNLP (Demos) , pages 90--98. Association for Computational Linguistics, 2023

  144. [169]

    Livebench: A challenging, contamination-limited LLM benchmark

    Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and M...

  145. [170]

    R e C a LL : Membership inference via relative conditional log-likelihoods

    Roy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang, Rong Ge, Jian Pei, Neil Zhenqiang Gong, and Bhuwan Dhingra. R e C a LL : Membership inference via relative conditional log-likelihoods. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Con...

  146. [171]

    Benchmark data contamination of large language models: A survey

    Cheng Xu, Shuhao Guan, Derek Greene, M Kechadi, et al. Benchmark data contamination of large language models: A survey. arXiv preprint arXiv:2406.04244, 2024

  147. [172]

    Rethinking benchmark and contamination for language models with rephrased samples

    Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E Gonzalez, and Ion Stoica. Rethinking benchmark and contamination for language models with rephrased samples. arXiv preprint arXiv:2311.04850, 2023

  148. [173]

    Enhanced membership inference attacks against machine learning models

    Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri. Enhanced membership inference attacks against machine learning models. In ACM SIGSAC Conference on Computer and Communications Security, pages 3093--3106, 2022

  149. [174]

    Membership inference on llms in the wild, 2026

    Jiatong Yi and Yanyang Li. Membership inference on llms in the wild, 2026. URL https://arxiv.org/abs/2601.11314

  150. [175]

    Detecting data contamination in LLM s via in-context learning

    Micha Zawalski, Meriem Boubdir, Klaudia Ba azy, Besmira Nushi, and Pablo Ribalta. Detecting data contamination in LLM s via in-context learning. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=YlpaaYxx4t

  151. [176]

    Fine-tuning can help detect pretraining data from large language models

    Hengxiang Zhang, Songxin Zhang, Bingyi Jing, and Hongxin Wei. Fine-tuning can help detect pretraining data from large language models. In International Conference on Learning Representations, 2025 a . URL https://openreview.net/forum?id=X8dzvdkQwO

  152. [177]

    P a C o ST : Paired confidence significance testing for benchmark contamination detection in large language models

    Huixuan Zhang, Yun Lin, and Xiaojun Wan. P a C o ST : Paired confidence significance testing for benchmark contamination detection in large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics...

  153. [178]

    Min-k\ large language models

    Jingyang Zhang, Jingwei Sun, Eric Yeats, Yang Ouyang, Martin Kuo, Jianyi Zhang, Hao Frank Yang, and Hai Li. Min-k\ large language models. In International Conference on Learning Representations, 2025 b . URL https://openreview.net/forum?id=ZGkfoufDaU

  154. [179]

    Pretraining data detection for large language models: A divergence-based calibration method

    Weichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. Pretraining data detection for large language models: A divergence-based calibration method. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conferen...

  155. [180]

    Mmlu-cf: A contamination-free multi-task language understanding benchmark

    Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui, Qinzheng Sun, Shaoguang Mao, Xin Zhang, Ying Xin, Qiufeng Yin, Scarlett Li, et al. Mmlu-cf: A contamination-free multi-task language understanding benchmark. arXiv preprint arXiv:2412.15194, 2024

  156. [181]

    Don't make your llm an evaluation benchmark cheater

    Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. Don't make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023

  157. [182]

    Dyval: Dynamic evaluation of large language models for reasoning tasks

    Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. Dyval: Dynamic evaluation of large language models for reasoning tasks. In The Twelfth International Conference on Learning Representations, 2024 a . URL https://openreview.net/forum?id=gjfOL9z5Xr

  158. [183]

    CLEAN -- EVAL : Clean evaluation on contaminated large language models

    Wenhong Zhu, Hongkun Hao, Zhiwei He, Yun-Ze Song, Jiao Yueyang, Yumeng Zhang, Hanxu Hu, Yiran Wei, Rui Wang, and Hongyuan Lu. CLEAN -- EVAL : Clean evaluation on contaminated large language models. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Associ...

Pith tools

Reviewed May 22, 2026 · model on record in the stance chip above.