Pith. sign in

REVIEW 2 cited by

InvBERT: Reconstructing Text from Contextualized Word Embeddings by inverting the BERT pipeline

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.10104 v2 pith:XO5SOW6F submitted 2021-09-21 cs.CL

classification cs.CL
keywords textcontextualizedbertembeddingsanalyticalcertaincopyrightdtfs
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Digital Humanities and Computational Literary Studies apply text mining methods to investigate literature. Such automated approaches enable quantitative studies on large corpora which would not be feasible by manual inspection alone. However, due to copyright restrictions, the availability of relevant digitized literary works is limited. Derived Text Formats (DTFs) have been proposed as a solution. Here, textual materials are transformed in such a way that copyright-critical features are removed, but that the use of certain analytical methods remains possible. Contextualized word embeddings produced by transformer-encoders (like BERT) are promising candidates for DTFs because they allow for state-of-the-art performance on various analytical tasks and, at first sight, do not disclose the original text. However, in this paper we demonstrate that under certain conditions the reconstruction of the original copyrighted text becomes feasible and its publication in the form of contextualized token representations is not safe. Our attempts to invert BERT suggest, that publishing the encoder as a black box together with the contextualized embeddings is critical, since it allows to generate data to train a decoder with a reconstruction accuracy sufficient to violate copyright laws.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeepInvert: Semi-Supervised Embedding Inversion Against Obfuscated Language Models

    cs.CR 2026-08 conditional novelty 7.0 of 10

    DeepInvert uses unlabeled obfuscated embeddings to train an inversion model that recovers up to 73.5% of original tokens against ObfusLM, versus 26.2% for the previous best attack.

  2. BeamClean: Language Aware Embedding Reconstruction

    cs.CR 2025-05 conditional novelty 6.0 of 10

    BeamClean recovers significantly more tokens and PII from Gaussian or Laplacian perturbed embeddings than nearest-neighbor attacks by combining noise-parameter estimation with a language-model prior.

Pith tools