REVIEW 2 major objections 2 minor 166 references
Provable Joint Decontamination for Benchmarking Multiple Large Language Models
T0 review · 2 major / 2 minor · reviewed 2026-05-22 · grok-4.3
Pith's one-line read A conformal procedure selects shared benchmarks for multiple LLMs while provably controlling the global contamination rate.
desk verdict JECS formalizes joint multi-model decontamination via max-p aggregation and right-tail envelope reconstruction for GCR control, but the envelope may lose conservativeness under realistic cross-model dependence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The reconstruction of a conservative envelope for the distribution of maximum p-values, which rescales the statistics to enable joint control across models.
What would settle it
A simulation or real audit in which the fraction of actually contaminated items among those selected by JECS exceeds the nominal global contamination rate target.
Extended reading notes
Core claim
JECS computes per-model conformal p-values, aggregates them by the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure to select a benchmark with provable global contamination rate control.
Load-bearing premise
The conservative envelope built from right-tail observations above a data-driven threshold continues to upper-bound the true max-p null distribution when the individual conformal p-values meet their stated assumptions.
Editorial extensions
If this is right
- Provides a single decontaminated benchmark that supports fair performance comparisons across all audited models.
- Maintains the target global contamination rate control even when separate per-model detections would produce inconsistent selections.
- Achieves higher statistical power for identifying clean items compared to applying the maximum p-value directly without the envelope adjustment.
- Applies to various LLMs and benchmarks while consistently respecting the contamination control target.
Reading between the lines
- Practitioners could adopt this for standardized evaluation suites shared among multiple model developers to reduce hidden contamination effects.
- The envelope approach might generalize to other joint multiple-testing problems where the maximum statistic is of interest.
- Testing the method on benchmarks with known partial contamination could reveal how sensitive the data-driven threshold is in practice.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalizes multi-model benchmark decontamination as a joint selection problem and proposes Joint Envelope Conformal Selection (JECS). JECS computes per-model conformal p-values, aggregates them via the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure to the rescaled values to select a benchmark with provable global contamination rate (GCR) control under stated assumptions. Experiments across models and benchmarks are reported to show higher power than the max-p baseline while maintaining target GCR.
Significance. If the GCR control holds under the procedure's assumptions, the result would be significant for LLM evaluation: it would provide the first joint conformal method that enables fair cross-model comparisons by producing a single decontaminated benchmark with theoretical guarantees, rather than model-specific selections. The use of conformal p-values and adaptive BH is a natural extension of existing single-model decontamination work.
major comments (2)
- The central claim of provable GCR control rests on the conservative envelope reconstruction step, but the manuscript provides no explicit derivation, no statement of the full set of assumptions (e.g., on the dependence structure among per-model p-values), and no quantitative experimental results or tables in the provided description. This makes the guarantee impossible to verify from the text and is load-bearing for the main contribution.
- The envelope reconstruction from right-tail observations above a data-driven threshold (method description) may fail to stochastically dominate the true null distribution of the maximum when per-model conformal p-values for the same item are positively dependent, which is realistic for shared benchmark items. The data-driven threshold correlates with the observed maxima, and no explicit modeling or robustness argument for this dependence is given; a counterexample or additional proof under dependence would be required to support the GCR claim.
minor comments (2)
- The abstract claims 'extensive experiments' with higher power and GCR control but reports no numerical values, tables, or specific benchmarks/models; adding these would improve clarity.
- Notation for the envelope and rescaled p-values should be defined more precisely with equation numbers to allow direct reference in the proof.
Simulated Author's Rebuttal
We thank the referee for the thoughtful and detailed report. The comments correctly identify that the theoretical guarantees are central to the contribution, and we will revise the manuscript to make the derivation, assumptions, and empirical verification fully explicit and self-contained.
read point-by-point responses
-
Referee: The central claim of provable GCR control rests on the conservative envelope reconstruction step, but the manuscript provides no explicit derivation, no statement of the full set of assumptions (e.g., on the dependence structure among per-model p-values), and no quantitative experimental results or tables in the provided description. This makes the guarantee impossible to verify from the text and is load-bearing for the main contribution.
Authors: We agree that the current presentation would benefit from greater explicitness. In the revised manuscript we will add a dedicated subsection (and appendix) containing the full derivation of the conservative envelope, the precise statement of all assumptions (including marginal validity of the per-model conformal p-values and the conditions under which the right-tail envelope stochastically dominates the null distribution of the maximum), and a self-contained proof of GCR control. We will also expand the experimental section with additional tables that report empirical GCR on both real benchmarks and controlled synthetic dependence scenarios. revision: yes
-
Referee: The envelope reconstruction from right-tail observations above a data-driven threshold (method description) may fail to stochastically dominate the true null distribution of the maximum when per-model conformal p-values for the same item are positively dependent, which is realistic for shared benchmark items. The data-driven threshold correlates with the observed maxima, and no explicit modeling or robustness argument for this dependence is given; a counterexample or additional proof under dependence would be required to support the GCR claim.
Authors: The concern is well-taken. Positive dependence among p-values for the same item is plausible and the data-driven threshold introduces correlation with the observed maxima. In the revision we will (i) explicitly state the working assumption of conditional independence across models given the item (or, alternatively, provide a worst-case bound), (ii) supply a short proof sketch showing that the envelope remains conservative under this assumption, and (iii) include a small simulation study that injects controlled positive dependence and verifies that the realized GCR stays below the nominal level. If the dependence is judged too strong for the guarantee to hold, we will clearly delineate the limitation in the revised text. revision: partial
Circularity Check
JECS derivation relies on standard conformal and BH theory with independent envelope construction
full rationale
The paper proposes JECS by computing per-model conformal p-values (standard), aggregating via per-item max, reconstructing a conservative envelope of the max-p null from right-tail observations above a data-driven threshold, and feeding rescaled values into adaptive BH for GCR control. This chain builds on established conformal inference and multiple-testing results without reducing the claimed guarantee to a tautology, self-definition, or fitted input renamed as prediction. The envelope step is presented as a novel but externally motivated construction under stated assumptions rather than presupposing the target GCR; no load-bearing premise collapses to self-citation or ansatz smuggling. The procedure remains self-contained against external benchmarks for its core components.
Assumptions & free parameters
assumptions (2)
- domain assumption Per-model conformal p-values are valid under the null of no contamination for each model.
- domain assumption The max-p null distribution can be conservatively reconstructed from right-tail observations above a data-driven threshold.
Cite this review
Pith. "Pith review of Provable Joint Decontamination for Benchmarking Multiple Large Language Models." pith.science (2026). https://pith.science/paper/75LO3J3K
@misc{pith2026260521543,
author = {Pith},
title = {Pith review of: Provable Joint Decontamination for Benchmarking Multiple Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/75LO3J3K}},
note = {Machine review of arXiv:2605.21543}
}
read the original abstract
Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performance can be inflated and cross-model comparisons become unreliable. A broad line of training-data detection work designs scores to quantify how strongly a model memorizes a given data point, but these score-based methods lack theoretical guarantees. Recent conformal approaches provide provable false-identification control for a single model; however, applying them separately to each model can produce model-specific benchmarks, undermining fair comparison across models. In this work, we formalize multi-model benchmark decontamination as a joint selection problem and propose Joint Envelope Conformal Selection (JECS), a conformal procedure that enables global contamination rate (GCR) control under stated assumptions. Specifically, JECS computes per-model conformal p-values, aggregates them by the per-item maximum, and reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold. By applying the adaptive Benjamini-Hochberg (BH) procedure to the envelope-rescaled values, we select a benchmark with provable GCR control. Extensive experiments across various models and benchmarks demonstrate that JECS achieves higher power than the max-p baseline while consistently maintaining the target GCR control.
Figures
Figures from the paper (5 more)
Lean theorems connected to this paper
-
IndisputableMonolith/Cost/FunctionalEquation.leanwashburn_uniqueness_aczel unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
JECS computes per-model conformal p-values, aggregates them by the per-item maximum, reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold, and applies the adaptive Benjamini-Hochberg procedure
What do these tags mean?
- matches
- The paper's claim is directly supported by a theorem in the formal canon.
- supports
- The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
- extends
- The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
- uses
- The paper appears to rely on the theorem as machinery.
- contradicts
- The paper's claim conflicts with a theorem or certificate in the canon.
- unclear
- Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.
Reference graph
Works this paper leans on
- [1]
-
[2]
IEEE Symposium on Security and Privacy , pages=
Membership inference attacks against machine learning models , author=. IEEE Symposium on Security and Privacy , pages=. 2017 , organization=
work page 2017
-
[3]
International Conference on Learning Representations , year=
How much of my dataset did you use? Quantitative Data Usage Inference in Machine Learning , author=. International Conference on Learning Representations , year=
-
[4]
International Conference on Learning Representations , year=
Quantifying Memorization Across Neural Language Models , author=. International Conference on Learning Representations , year=
-
[5]
Regulation (EU) 2016/679 (General Data Protection Regulation) , howpublished =. 2016 , month =
work page 2016
-
[6]
The California Consumer Privacy Act of 2018 (CCPA) , institution =. 2018 , number =
work page 2018
-
[7]
IEEE Conference on Secure and Trustworthy Machine Learning , pages=
Position: Membership Inference Attacks Cannot Prove That a Model was Trained on Your Data , author=. IEEE Conference on Secure and Trustworthy Machine Learning , pages=. 2025 , organization=
work page 2025
-
[8]
arXiv preprint arXiv:2309.10677 , year=
Estimating contamination via perplexity: Quantifying memorisation in language model evaluation , author=. arXiv preprint arXiv:2309.10677 , year=
Show all 166 references
-
[9]
International Conference on Learning Representations , year=
Detecting Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=
-
[10]
IEEE Symposium on Security and Privacy (SP) , pages=
Membership inference attacks from first principles , author=. IEEE Symposium on Security and Privacy (SP) , pages=. 2022 , organization=
2022
-
[11]
2018 , organization=
Privacy risk in machine learning: Analyzing the connection to overfitting , author=. 2018 , organization=
2018
-
[12]
Ahmed Salem and Yang Zhang and Mathias Humbert and Pascal Berrang and Mario Fritz and Michael Backes , title =
-
[13]
ACM SIGSAC Conference on Computer and Communications Security , pages=
Enhanced membership inference attacks against machine learning models , author=. ACM SIGSAC Conference on Computer and Communications Security , pages=
-
[14]
Advances in Neural Information Processing Systems , volume=
Scalable membership inference attacks via quantile regression , author=. Advances in Neural Information Processing Systems , volume=
-
[15]
The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
Membership Inference Attacks against Large Vision-Language Models , author=. The Thirty-eighth Annual Conference on Neural Information Processing Systems , year=
-
[16]
International Conference on Machine Learning , year=
Low-Cost High-Power Membership Inference Attacks , author=. International Conference on Machine Learning , year=
-
[17]
Advances in Neural Information Processing Systems , year=
Membership Inference Attacks against Fine-tuned Large Language Models via Self-prompt Calibration , author=. Advances in Neural Information Processing Systems , year=
-
[18]
USENIX Security Symposium , pages=
Extracting training data from large language models , author=. USENIX Security Symposium , pages=
-
[19]
arXiv preprint arXiv:2503.07482 , year=
Efficient Membership Inference Attacks by Bayesian Neural Network , author=. arXiv preprint arXiv:2503.07482 , year=
-
[20]
International Conference on Learning Representations , year=
Fine-tuning can Help Detect Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=
-
[21]
Journal of Machine Learning Research , volume=
Selection by prediction with conformal p-values , author=. Journal of Machine Learning Research , volume=
-
[22]
Journal of the Royal Statistical Society , volume=
Controlling the false discovery rate: a practical and powerful approach to multiple testing , author=. Journal of the Royal Statistical Society , volume=. 1995 , publisher=
1995
-
[23]
Annals of Statistics , pages=
The control of the false discovery rate in multiple testing under dependency , author=. Annals of Statistics , pages=. 2001 , publisher=
2001
-
[24]
International Conference on Learning Representations , year=
Proving test set contamination in black-box language models , author=. International Conference on Learning Representations , year=
-
[25]
International Conference on Learning Representations , year=
A Statistical Approach for Controlled Training Data Detection , author=. International Conference on Learning Representations , year=
-
[26]
Deyao Zhu and Jun Chen and Xiaoqian Shen and Xiang Li and Mohamed Elhoseiny , booktitle=. Mini. 2024 , url=
2024
-
[27]
Advances in Neural Information Processing Systems , volume=
Visual instruction tuning , author=. Advances in Neural Information Processing Systems , volume=
-
[29]
International Conference on Machine Learning , pages=
Pythia: A suite for analyzing large language models across training and scaling , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[30]
arXiv preprint arXiv:2205.01068 , year=
Opt: Open pre-trained transformer language models , author=. arXiv preprint arXiv:2205.01068 , year=
-
[32]
IEEE Conference on Computer Vision and Pattern Recognition , pages=
Deep residual learning for image recognition , author=. IEEE Conference on Computer Vision and Pattern Recognition , pages=
-
[33]
IEEE Conference on Computer Vision and Pattern Recognition , pages=
Densely connected convolutional networks , author=. IEEE Conference on Computer Vision and Pattern Recognition , pages=
-
[34]
International Conference on Machine Learning , year=
Andr. International Conference on Machine Learning , year=
-
[35]
2019 , url=
Language Models are Unsupervised Multitask Learners , author=. 2019 , url=
2019
-
[36]
arXiv preprint arXiv:2101.00027 , year=
The Pile: An 800GB Dataset of Diverse Text for Language Modeling , author=. arXiv preprint arXiv:2101.00027 , year=
-
[37]
and Lapata, Mirella
Narayan, Shashi and Cohen, Shay B. and Lapata, Mirella. Don ' t Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme Summarization. Conference on Empirical Methods in Natural Language Processing. doi:10.18653/v1/D18-1206
-
[38]
Yucheng Li and Frank Guerin and Chenghua Lin , title =
-
[39]
2009 , journal=
Learning multiple layers of features from tiny images , author=. 2009 , journal=
2009
-
[40]
USENIX Security Symposium , pages=
Systematic evaluation of privacy risks of machine learning models , author=. USENIX Security Symposium , pages=
-
[41]
Berkeley Symposium on Mathematical Statistics and Probability , volume=
On measures of entropy and information , author=. Berkeley Symposium on Mathematical Statistics and Probability , volume=. 1961 , organization=
1961
-
[42]
Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models
Dong, Yihong and Jiang, Xue and Liu, Huanyu and Jin, Zhi and Gu, Bin and Yang, Mengfei and Li, Ge. Generalization or Memorization: Data Contamination and Trustworthy Evaluation for Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024. 2024...
2024 doi
-
[45]
doi:10.5281/zenodo.12608602 , url =
Gao, Leo and Tow, Jonathan and Abbasi, Baber and Biderman, Stella and Black, Sid and DiPofi, Anthony and Foster, Charles and Golding, Laurence and Hsu, Jeffrey and Le Noac'h, Alain and Li, Haonan and McDonell, Kyle and Muennighoff, Niklas and Ociepa, Chris and Phang, Jason and...
-
[46]
NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark
Sainz, Oscar and Campos, Jon and Garc \'i a-Ferrero, Iker and Etxaniz, Julen and de Lacalle, Oier Lopez and Agirre, Eneko. NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. Findings of the Association for Computational Linguistics: EM...
2023 doi
-
[47]
Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLM s
Balloccu, Simone and Schmidtov \'a , Patr \'i cia and Lango, Mateusz and Dusek, Ondrej. Leak, Cheat, Repeat: Data Contamination and Evaluation Malpractices in Closed-Source LLM s. Proceedings of the 18th Conference of the European Chapter of the Association for Computational L...
2024 doi
-
[48]
Proceedings of the IEEE Symposium on Security and Privacy (SP) , pages=
Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning , author=. Proceedings of the IEEE Symposium on Security and Privacy (SP) , pages=. 2019 , organization=
2019
-
[49]
International Conference on Machine Learning , pages=
White-box vs black-box: Bayes optimal strategies for membership inference , author=. International Conference on Machine Learning , pages=. 2019 , organization=
2019
-
[50]
Advances in Neural Information Processing Systems , year =
Gaussian Membership Inference Privacy , author =. Advances in Neural Information Processing Systems , year =
-
[51]
IEEE Transactions on Computational Social Systems , volume=
Socinf: Membership inference attacks on social media health data with machine learning , author=. IEEE Transactions on Computational Social Systems , volume=. 2019 , publisher=
2019
-
[52]
Lauren Watson and Chuan Guo and Graham Cormode and Alexandre Sablayrolles , title =
-
[54]
Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =
Inbal Magar and Roy Schwartz , title =. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics , pages =
-
[55]
International Conference on Learning Representations , year=
Min-K\ author=. International Conference on Learning Representations , year=
-
[56]
2024 , booktitle=
Do Membership Inference Attacks Work on Large Language Models? , author=. 2024 , booktitle=
2024
-
[58]
International Conference on Learning Representations , year=
Infilling Score: A Pretraining Data Detection Algorithm for Large Language Models , author=. International Conference on Learning Representations , year=
-
[60]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
A direct approach to false discovery rates , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2002 , publisher=
2002
-
[61]
Vladimir Vovk and Ilia Nouretdinov and Alex Gammerman , title =
-
[62]
2005 , publisher=
Algorithmic learning in a random world , author=. 2005 , publisher=
2005
-
[63]
The Annals of Statistics , volume=
Testing for outliers with conformal p-values , author=. The Annals of Statistics , volume=. 2023 , publisher=
2023
-
[64]
Computational Statistics & Data Analysis , volume=
Beta kernel estimators for density functions , author=. Computational Statistics & Data Analysis , volume=. 1999 , publisher=
1999
-
[65]
Journal of Nonparametric Statistics , volume=
Bias reductions for beta kernel estimation , author=. Journal of Nonparametric Statistics , volume=. 2016 , publisher=
2016
-
[66]
Journal of the American Statistical Association , volume=
Improvement of kernel type density estimators , author=. Journal of the American Statistical Association , volume=. 1977 , publisher=
1977
-
[67]
Pierre Neuvial , title =. J. Mach. Learn. Res. , volume =
-
[68]
Biometrika , volume=
Adaptive linear step-up procedures that control the false discovery rate , author=. Biometrika , volume=. 2006 , publisher=
2006
-
[69]
Journal of Educational and Behavioral Statistics , volume=
On the adaptive control of the false discovery rate in multiple testing with independent statistics , author=. Journal of Educational and Behavioral Statistics , volume=. 2000 , publisher=
2000
-
[70]
Proceedings of the Sixteenth International Conference on Machine Learning , pages =
Vovk, Volodya and Gammerman, Alexander and Saunders, Craig , title =. Proceedings of the Sixteenth International Conference on Machine Learning , pages =. 1999 , isbn =
1999
-
[71]
ACM transactions on intelligent systems and technology , volume=
A survey on evaluation of large language models , author=. ACM transactions on intelligent systems and technology , volume=. 2024 , publisher=
2024
-
[72]
Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
Investigating data contamination in modern benchmarks for large language models , author=. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , pages=
2024
-
[74]
Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark , author=. Findings of the Association for Computational Linguistics: EMNLP 2023 , pages=
2023
-
[76]
Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source LLMs , author=. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[78]
arXiv preprint arXiv:2211.15533 , year=
The stack: 3 tb of permissively licensed source code , author=. arXiv preprint arXiv:2211.15533 , year=
-
[79]
arXiv preprint arXiv:2305.06161 , year=
Starcoder: may the source be with you! , author=. arXiv preprint arXiv:2305.06161 , year=
-
[80]
Advances in Neural Information Processing Systems , volume=
Redpajama: an open dataset for training large language models , author=. Advances in Neural Information Processing Systems , volume=
-
[82]
Advances in Neural Information Processing Systems , volume=
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only , author=. Advances in Neural Information Processing Systems , volume=
-
[83]
2024 , month = jul, day =
2024
-
[84]
Forty-third International Conference on Machine Learning , year=
Provable Training Data Identification for Large Language Models , author=. Forty-third International Conference on Machine Learning , year=
-
[85]
The Fourteenth International Conference on Learning Representations , year=
Multi-Condition Conformal Selection , author=. The Fourteenth International Conference on Learning Representations , year=
-
[86]
Forty-second International Conference on Machine Learning , year=
Multivariate Conformal Selection , author=. Forty-second International Conference on Machine Learning , year=
-
[87]
2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=
Sok: Membership inference attacks on llms are rushing nowhere (and how to fix it) , author=. 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML) , pages=. 2025 , organization=
2025
-
[88]
Advances in Neural Information Processing Systems , volume=
LLM Dataset Inference: Did you train on my dataset? , author=. Advances in Neural Information Processing Systems , volume=
-
[89]
Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=
Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach , author=. Journal of the Royal Statistical Society Series B: Statistical Methodology , volume=. 2004 , publisher=
2004
-
[90]
Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Detecting Non-Membership in LLM Training Data via Rank Correlations , author=. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
-
[91]
Forty-second International Conference on Machine Learning , year=
How Contaminated Is Your Benchmark? Measuring Dataset Leakage in Large Language Models with Kernel Divergence , author=. Forty-second International Conference on Machine Learning , year=
-
[92]
The Fourteenth International Conference on Learning Representations , year=
BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language Models , author=. The Fourteenth International Conference on Learning Representations , year=
-
[94]
The Twelfth International Conference on Learning Representations , year=
Proving Test Set Contamination in Black-Box Language Models , author=. The Twelfth International Conference on Learning Representations , year=
-
[97]
2026 , url=
Yihao LIU and Xinqi LYU and Dong Wang and Yanjie Li and Bin Xiao , booktitle=. 2026 , url=
2026
-
[98]
Yuke Hu and Zheng Li and Zhihao Liu and Yang Zhang and Zhan Qin and Kui Ren and Chun Chen , title =
-
[99]
Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
Practical membership inference attacks against large-scale multi-modal models: A pilot study , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=
-
[101]
2026 , eprint=
Membership Inference on LLMs in the Wild , author=. 2026 , eprint=
2026
-
[102]
Detecting Data Contamination in
Micha. Detecting Data Contamination in. The Fourteenth International Conference on Learning Representations , year=
-
[104]
The Twelfth International Conference on Learning Representations , year=
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks , author=. The Twelfth International Conference on Learning Representations , year=
-
[106]
Koala: An Index for Quantifying Overlaps with Pre-training Corpora , booktitle =
Thuy. Koala: An Index for Quantifying Overlaps with Pre-training Corpora , booktitle =
-
[107]
GitHub repository , howpublished =
Yanyang Li , title =. GitHub repository , howpublished =. 2024 , publisher =
2024
-
[108]
LiveBench: A Challenging, Contamination-Limited
Colin White and Samuel Dooley and Manley Roberts and Arka Pal and Benjamin Feuer and Siddhartha Jain and Ravid Shwartz-Ziv and Neel Jain and Khalid Saifullah and Sreemanti Dey and Shubh-Agrawal and Sandeep Singh Sandha and Siddartha Venkat Naidu and Chinmay Hegde and Yann LeCu...
2025
-
[109]
2024 , eprint=
C ^2 LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation , author=. 2024 , eprint=
2024
-
[110]
Sebastian Bordt and Suraj Srinivas and Valentyn Boreiko and Ulrike von Luxburg , title =
-
[111]
2008 , publisher=
Inductive conformal prediction: Theory and application to neural networks , author=. 2008 , publisher=
2008
-
[112]
Biometrika , pages=
Model-free selective inference under covariate shift via weighted conformal p-values , author=. Biometrika , pages=. 2025 , publisher=
2025
-
[113]
Advances in Neural Information Processing Systems , volume=
Conformal alignment: Knowing when to trust foundation models with guarantees , author=. Advances in Neural Information Processing Systems , volume=
-
[114]
Journal of Chemical Information and Modeling , volume=
Conformal selection for efficient and accurate compound screening in drug discovery , author=. Journal of Chemical Information and Modeling , volume=. 2025 , publisher=
2025
-
[115]
Time Travel in
Shahriar Golchin and Mihai Surdeanu , booktitle=. Time Travel in. 2024 , url=
2024
-
[116]
International Conference on Learning Representations , year=
Fine-tuning Can Help Detect Pretraining Data from Large Language Models , author=. International Conference on Learning Representations , year=
-
[117]
Gpt-4 technical report
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023
2023 arXiv
-
[118]
Conformal selection for efficient and accurate compound screening in drug discovery
Tian Bai, Peng Tang, Yuting Xu, Vladimir Svetnik, Bingjia Yang, Abbas Khalili, Xiang Yu, and Archer Y Yang. Conformal selection for efficient and accurate compound screening in drug discovery. Journal of Chemical Information and Modeling, 65 0 (24): 0 13070--13085, 2025 a
2025
-
[119]
Tian Bai, Yue Zhao, Xiang Yu, and Archer Y. Yang. Multivariate conformal selection. In Forty-second International Conference on Machine Learning, 2025 b . URL https://openreview.net/forum?id=g2tr7nA4pS
2025
-
[120]
Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms
Simone Balloccu, Patr \' cia Schmidtov \'a , Mateusz Lango, and Ond r ej Du s ek. Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Lingu...
2024
-
[121]
Testing for outliers with conformal p-values
Stephen Bates, Emmanuel Cand \`e s, Lihua Lei, Yaniv Romano, and Matteo Sesia. Testing for outliers with conformal p-values. The Annals of Statistics, 51 0 (1): 0 149--178, 2023
2023
-
[122]
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society, 57 0 (1): 0 289--300, 1995
1995
-
[123]
The control of the false discovery rate in multiple testing under dependency
Yoav Benjamini and Daniel Yekutieli. The control of the false discovery rate in multiple testing under dependency. Annals of Statistics, pages 1165--1188, 2001
2001
-
[124]
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In Intern...
2023
-
[125]
GPT - N eo X -20 B : An open-source autoregressive language model
Sidney Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, Usvsn Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. GPT - N eo X -20 ...
2022 doi
-
[126]
How much can we forget about data contamination? In ICML , Proceedings of Machine Learning Research
Sebastian Bordt, Suraj Srinivas, Valentyn Boreiko, and Ulrike von Luxburg. How much can we forget about data contamination? In ICML , Proceedings of Machine Learning Research. PMLR / OpenReview.net, 2025
2025
-
[127]
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. Extracting training data from large language models. In USENIX Security Symposium, pages 2633--2650, 2021
2021
-
[128]
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. A survey on evaluation of large language models. ACM transactions on intelligent systems and technology, 15 0 (3): 0 1--45, 2024
2024
-
[129]
A survey on data contamination for large language models
Yuxing Cheng, Yi Chang, and Yuan Wu. A survey on data contamination for large language models. arXiv preprint arXiv:2502.14425, 2025
2025
-
[130]
How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence
Hyeong Kyu Choi, Maxim Khanov, Hongxin Wei, and Yixuan Li. How contaminated is your benchmark? measuring dataset leakage in large language models with kernel divergence. In Forty-second International Conference on Machine Learning, 2025. URL https://openreview.net/forum?id=wVDR2qmE28
2025
-
[131]
Investigating data contamination in modern benchmarks for large language models
Chunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, and Arman Cohan. Investigating data contamination in modern benchmarks for large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human...
2024
-
[132]
Do membership inference attacks work on large language models? In Conference on Language Modeling, 2024
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi. Do membership inference attacks work on large language models? In Conference on Language Modeling, 2024
2024
-
[133]
Oliveira, and Lei Li
Andr \'e Vicente Duarte, Xuandong Zhao, Arlindo L. Oliveira, and Lei Li. DE - COP : Detecting copyrighted content in language models training data. In International Conference on Machine Learning, 2024. URL https://openreview.net/forum?id=LO4xhXmFal
2024
-
[134]
Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act)
European Parliament and Council of the European Union . Regulation (EU) 2024/1689 of the European Parliament and of the Council laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) . Official Journal of the European Union, OJ L, 2024/1689, 12 J...
2024
-
[135]
Time travel in LLM s: Tracing data contamination in large language models
Shahriar Golchin and Mihai Surdeanu. Time travel in LLM s: Tracing data contamination in large language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=2Rwq6c3tvr
2024
-
[136]
Conformal alignment: Knowing when to trust foundation models with guarantees
Yu Gui, Ying Jin, and Zhimei Ren. Conformal alignment: Knowing when to trust foundation models with guarantees. Advances in Neural Information Processing Systems, 37: 0 73884--73919, 2024
2024
-
[137]
Multi-condition conformal selection
Qingyang Hao, Wenbo Liao, Bingyi Jing, and Hongxin Wei. Multi-condition conformal selection. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=giL8Q1V26J
2026
-
[138]
Membership inference attacks against vision-language models
Yuke Hu, Zheng Li, Zhihao Liu, Yang Zhang, Zhan Qin, Kui Ren, and Chun Chen. Membership inference attacks against vision-language models. In USENIX Security Symposium , pages 1589--1608. USENIX Association, 2025 a
2025
-
[139]
A statistical approach for controlled training data detection
Zirui Hu, Yingjie Wang, Zheng Zhang, Hong Chen, and Dacheng Tao. A statistical approach for controlled training data detection. In International Conference on Learning Representations, 2025 b . URL https://openreview.net/forum?id=XAN8G0rvoB
2025
-
[140]
Selection by prediction with conformal p-values
Ying Jin and Emmanuel J Cand \`e s. Selection by prediction with conformal p-values. Journal of Machine Learning Research, 24 0 (244): 0 1--41, 2023
2023
-
[141]
Model-free selective inference under covariate shift via weighted conformal p-values
Ying Jin and Emmanuel J Cand \`e s. Model-free selective inference under covariate shift via weighted conformal p-values. Biometrika, page asaf066, 2025
2025
-
[142]
Sampling-based pseudo-likelihood for membership inference attacks
Masahiro Kaneko, Youmi Ma, Yuki Wata, and Naoaki Okazaki. Sampling-based pseudo-likelihood for membership inference attacks. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL ...
2025 doi
-
[143]
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models. arXiv preprint arXiv:2001.08361, 2020
2001 arXiv
-
[144]
Practical membership inference attacks against large-scale multi-modal models: A pilot study
Myeongseob Ko, Ming Jin, Chenguang Wang, and Ruoxi Jia. Practical membership inference attacks against large-scale multi-modal models: A pilot study. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 4871--4881, 2023
2023
-
[145]
Task contamination: language models may not be few-shot anymore
Changmao Li and Jeffrey Flanigan. Task contamination: language models may not be few-shot anymore. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Sy...
2024 doi
-
[146]
Awesome data contamination
Yanyang Li. Awesome data contamination. https://github.com/lyy1994/awesome-data-contamination, 2024
2024
-
[147]
Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction
Yucheng Li, Frank Guerin, and Chenghua Lin. Latesteval: Addressing data contamination in language model evaluation through dynamic and time-sensitive test construction. In AAAI Conference on Artificial Intelligence , pages 18600--18607. AAAI Press, 2024 a
2024
-
[148]
An open-source data contamination report for large language models
Yucheng Li, Yunhao Guo, Frank Guerin, and Chenghua Lin. An open-source data contamination report for large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 528--541, Mia...
2024 doi
-
[149]
Membership inference attacks against large vision-language models
Zhan Li, Yongtao Wu, Yihang Chen, Francesco Tonin, Elias Abad Rocamora, and Volkan Cevher. Membership inference attacks against large vision-language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024 c . URL https://openreview.net/fo...
2024
-
[150]
LOMIA : Label-only membership inference attacks against pre-trained large vision-language models
Yihao LIU, Xinqi LYU, Dong Wang, Yanjie Li, and Bin Xiao. LOMIA : Label-only membership inference attacks against pre-trained large vision-language models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2026. URL https://openreview.net/forum?id...
2026
-
[151]
Provable training data identification for large language models
Zhenlong Liu, Hao Zeng, Weiran Huang, and Hongxin Wei. Provable training data identification for large language models. In Forty-third International Conference on Machine Learning, 2026. URL https://arxiv.org/abs/2510.09717. arXiv preprint arXiv:2510.09717
2026
-
[152]
Llm dataset inference: Did you train on my dataset? Advances in Neural Information Processing Systems, 37: 0 124069--124092, 2024
Pratyush Maini, Hengrui Jia, Nicolas Papernot, and Adam Dziedzic. Llm dataset inference: Did you train on my dataset? Advances in Neural Information Processing Systems, 37: 0 124069--124092, 2024
2024
-
[153]
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. Membership inference attacks against language models via neighbourhood comparison. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findi...
2023 doi
-
[154]
Chatterji, Faisal Ladhak, and Tatsunori Hashimoto
Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto. Proving test set contamination in black-box language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=KS8mIvetg2
2024
-
[155]
Inductive conformal prediction: Theory and application to neural networks
Harris Papadopoulos. Inductive conformal prediction: Theory and application to neural networks. INTECH Open Access Publisher Rijeka, 2008
2008
-
[156]
The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. The refinedweb dataset for falcon llm: Outperforming curated corpora with web data only. Advances in Neural Inf...
2023
-
[157]
Infilling score: A pretraining data detection algorithm for large language models
Negin Raoof, Litu Rout, Giannis Daras, Sujay Sanghavi, Constantine Caramanis, Sanjay Shakkottai, and Alex Dimakis. Infilling score: A pretraining data detection algorithm for large language models. In International Conference on Learning Representations, 2025. URL https://open...
2025
-
[158]
Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark
Oscar Sainz, Jon Campos, Iker Garc \' a-Ferrero, Julen Etxaniz, Oier Lopez de Lacalle, and Eneko Agirre. Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages...
2023
-
[159]
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer. Detecting pretraining data from large language models. In International Conference on Learning Representations, 2024. URL https://openreview.net/forum?id=zWqr3MQuNs
2024
-
[160]
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov. Membership inference attacks against machine learning models. In IEEE Symposium on Security and Privacy, pages 3--18. IEEE, 2017
2017
-
[161]
Systematic evaluation of privacy risks of machine learning models
Liwei Song and Prateek Mittal. Systematic evaluation of privacy risks of machine learning models. In USENIX Security Symposium, pages 2615--2632, 2021
2021
-
[162]
Beyondbench: Contamination-resistant evaluation of reasoning in language models
Gaurav Srivastava, Aafiya Shamshad Hussain, Zhenyu Bi, Swastik Roy, Priya Pitre, Meng Lu, Morteza Ziyadi, and Xuan Wang. Beyondbench: Contamination-resistant evaluation of reasoning in language models. In The Fourteenth International Conference on Learning Representations, 202...
2026
-
[163]
A direct approach to false discovery rates
John D Storey. A direct approach to false discovery rates. Journal of the Royal Statistical Society Series B: Statistical Methodology, 64 0 (3): 0 479--498, 2002
2002
-
[164]
Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach
John D Storey, Jonathan E Taylor, and David Siegmund. Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach. Journal of the Royal Statistical Society Series B: Statistical Methodology, 66 0 (1): 0 1...
2004
-
[165]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth \'e e Lacroix, Baptiste Rozi \`e re, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023
2023 arXiv
-
[166]
Testing exchangeability on-line
Vladimir Vovk, Ilia Nouretdinov, and Alex Gammerman. Testing exchangeability on-line. In International Conference on Machine Learning , pages 768--775. AAAI Press, 2003
2003
-
[167]
Algorithmic learning in a random world
Vladimir Vovk, Alexander Gammerman, and Glenn Shafer. Algorithmic learning in a random world. Springer, 2005
2005
-
[168]
Koala: An index for quantifying overlaps with pre-training corpora
Thuy - Trang Vu, Xuanli He, Gholamreza Haffari, and Ehsan Shareghi. Koala: An index for quantifying overlaps with pre-training corpora. In EMNLP (Demos) , pages 90--98. Association for Computational Linguistics, 2023
2023
-
[169]
Livebench: A challenging, contamination-limited LLM benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and M...
2025
-
[170]
R e C a LL : Membership inference via relative conditional log-likelihoods
Roy Xie, Junlin Wang, Ruomin Huang, Minxing Zhang, Rong Ge, Jian Pei, Neil Zhenqiang Gong, and Bhuwan Dhingra. R e C a LL : Membership inference via relative conditional log-likelihoods. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Con...
2024 doi
-
[171]
Benchmark data contamination of large language models: A survey
Cheng Xu, Shuhao Guan, Derek Greene, M Kechadi, et al. Benchmark data contamination of large language models: A survey. arXiv preprint arXiv:2406.04244, 2024
2024
-
[172]
Rethinking benchmark and contamination for language models with rephrased samples
Shuo Yang, Wei-Lin Chiang, Lianmin Zheng, Joseph E Gonzalez, and Ion Stoica. Rethinking benchmark and contamination for language models with rephrased samples. arXiv preprint arXiv:2311.04850, 2023
2023
-
[173]
Enhanced membership inference attacks against machine learning models
Jiayuan Ye, Aadyaa Maddi, Sasi Kumar Murakonda, Vincent Bindschaedler, and Reza Shokri. Enhanced membership inference attacks against machine learning models. In ACM SIGSAC Conference on Computer and Communications Security, pages 3093--3106, 2022
2022
-
[174]
Membership inference on llms in the wild, 2026
Jiatong Yi and Yanyang Li. Membership inference on llms in the wild, 2026. URL https://arxiv.org/abs/2601.11314
2026
-
[175]
Detecting data contamination in LLM s via in-context learning
Micha Zawalski, Meriem Boubdir, Klaudia Ba azy, Besmira Nushi, and Pablo Ribalta. Detecting data contamination in LLM s via in-context learning. In The Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=YlpaaYxx4t
2026
-
[176]
Fine-tuning can help detect pretraining data from large language models
Hengxiang Zhang, Songxin Zhang, Bingyi Jing, and Hongxin Wei. Fine-tuning can help detect pretraining data from large language models. In International Conference on Learning Representations, 2025 a . URL https://openreview.net/forum?id=X8dzvdkQwO
2025
-
[177]
P a C o ST : Paired confidence significance testing for benchmark contamination detection in large language models
Huixuan Zhang, Yun Lin, and Xiaojun Wan. P a C o ST : Paired confidence significance testing for benchmark contamination detection in large language models. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics...
2024 doi
-
[178]
Min-k\ large language models
Jingyang Zhang, Jingwei Sun, Eric Yeats, Yang Ouyang, Martin Kuo, Jianyi Zhang, Hao Frank Yang, and Hai Li. Min-k\ large language models. In International Conference on Learning Representations, 2025 b . URL https://openreview.net/forum?id=ZGkfoufDaU
2025
-
[179]
Pretraining data detection for large language models: A divergence-based calibration method
Weichao Zhang, Ruqing Zhang, Jiafeng Guo, Maarten de Rijke, Yixing Fan, and Xueqi Cheng. Pretraining data detection for large language models: A divergence-based calibration method. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Proceedings of the 2024 Conferen...
2024 doi
-
[180]
Mmlu-cf: A contamination-free multi-task language understanding benchmark
Qihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui, Qinzheng Sun, Shaoguang Mao, Xin Zhang, Ying Xin, Qiufeng Yin, Scarlett Li, et al. Mmlu-cf: A contamination-free multi-task language understanding benchmark. arXiv preprint arXiv:2412.15194, 2024
2024
-
[181]
Don't make your llm an evaluation benchmark cheater
Kun Zhou, Yutao Zhu, Zhipeng Chen, Wentong Chen, Wayne Xin Zhao, Xu Chen, Yankai Lin, Ji-Rong Wen, and Jiawei Han. Don't make your llm an evaluation benchmark cheater. arXiv preprint arXiv:2311.01964, 2023
2023
-
[182]
Dyval: Dynamic evaluation of large language models for reasoning tasks
Kaijie Zhu, Jiaao Chen, Jindong Wang, Neil Zhenqiang Gong, Diyi Yang, and Xing Xie. Dyval: Dynamic evaluation of large language models for reasoning tasks. In The Twelfth International Conference on Learning Representations, 2024 a . URL https://openreview.net/forum?id=gjfOL9z5Xr
2024
-
[183]
CLEAN -- EVAL : Clean evaluation on contaminated large language models
Wenhong Zhu, Hongkun Hao, Zhiwei He, Yun-Ze Song, Jiao Yueyang, Yumeng Zhang, Hanxu Hu, Yiran Wei, Rui Wang, and Hongyuan Lu. CLEAN -- EVAL : Clean evaluation on contaminated large language models. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Associ...
2024 doi
Reviewed May 22, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.