Pith. sign in

Data contamination: From memorization to exploitation

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 2 2023 1

roles

background 1

polarities

support 1

representative citing papers

Detecting Pretraining Data from Large Language Models

cs.CL · 2023-10-25 · conditional · novelty 7.0

Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.

Dissecting model behavior through agent trajectories

cs.AI · 2026-06-16 · unverdicted · novelty 5.0

SSA harness matches frontier model pass@1 scores on agent benchmarks and 138k trajectory analysis in code state-spaces shows model-specific differences in edit frequency, testing activity, and phase transitions.

citing papers explorer

Showing 3 of 3 citing papers.

  • Detecting Pretraining Data from Large Language Models cs.CL · 2023-10-25 · conditional · none · ref 36

    Min-K% Prob detects pretraining data in LLMs by flagging outlier low-probability words in text, achieving 7.4% better performance than prior methods on the new WIKIMIA benchmark.

  • Dissecting model behavior through agent trajectories cs.AI · 2026-06-16 · unverdicted · none · ref 28

    SSA harness matches frontier model pass@1 scores on agent benchmarks and 138k trajectory analysis in code state-spaces shows model-specific differences in edit frequency, testing activity, and phase transitions.

  • DualEval: Joint Model-Item Calibration for Unified LLM Evaluation cs.LG · 2026-06-24 · unverdicted · none · ref 2

    DualEval jointly calibrates LLM abilities and item difficulties/sharpness in a shared latent space using static labels and reward-model scores to unify benchmark and arena-style evaluation.