Pith. sign in

Analyzing Information Leakage of Updates to Natural Language Models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

To continuously improve quality and reflect changes in data, machine learning applications have to regularly retrain and update their core models. We show that a differential analysis of language model snapshots before and after an update can reveal a surprising amount of detailed information about changes in the training data. We propose two new metrics---\emph{differential score} and \emph{differential rank}---for analyzing the leakage due to updates of natural language models. We perform leakage analysis using these metrics across models trained on several different datasets using different methods and configurations. We discuss the privacy implications of our findings, propose mitigation strategies and evaluate their effect.

citation-role summary

background 1

citation-polarity summary

fields

cs.CY 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

AI Safety for Everyone

cs.CY · 2025-02-13 · conditional · novelty 5.0

A systematic review of 383 papers argues that AI safety research already covers a wide spectrum of concrete, near-term concerns and should be understood as part of traditional technological safety practice.

citing papers explorer

Showing 1 of 1 citing paper.

  • AI Safety for Everyone cs.CY · 2025-02-13 · conditional · none · ref 89 · internal anchor

    A systematic review of 383 papers argues that AI safety research already covers a wide spectrum of concrete, near-term concerns and should be understood as part of traditional technological safety practice.