REVIEW 2 cited by
Neural Text Sanitization with Privacy Risk Indicators: An Empirical Analysis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Text sanitization is the task of redacting a document to mask all occurrences of (direct or indirect) personal identifiers, with the goal of concealing the identity of the individual(s) referred in it. In this paper, we consider a two-step approach to text sanitization and provide a detailed analysis of its empirical performance on two recently published datasets: the Text Anonymization Benchmark (Pil\'an et al., 2022) and a collection of Wikipedia biographies (Papadopoulou et al., 2022). The text sanitization process starts with a privacy-oriented entity recognizer that seeks to determine the text spans expressing identifiable personal information. This privacy-oriented entity recognizer is trained by combining a standard named entity recognition model with a gazetteer populated by person-related terms extracted from Wikidata. The second step of the text sanitization process consists in assessing the privacy risk associated with each detected text span, either isolated or in combination with other text spans. We present five distinct indicators of the re-identification risk, respectively based on language model probabilities, text span classification, sequence labelling, perturbations, and web search. We provide a contrastive analysis of each privacy indicator and highlight their benefits and limitations, notably in relation to the available labeled data.
Forward citations
Cited by 2 Pith papers
-
FindMyText: Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora
FindMyText detects near-verbatim text containment via fingerprint chains and outperforms shared-fingerprint, BM25 and dense-retrieval baselines on Wikipedia, ArXiv and HPLT web data.
-
Preempting Text Sanitization Utility in Resource-Constrained Privacy-Preserving LLM Interactions
A local small language model can predict when a differentially private sanitized prompt will still yield useful LLM output, saving up to 20% of wasted API calls, and an exact-nearest-neighbor implementation of the dX-...
Discussion (0). Continue with ORCID to comment.