Pith. sign in

REVIEW 2 cited by

A Public and Reproducible Assessment of the Topics API on Real Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.19577 v3 pith:CJVEFHXE submitted 2024-03-28 cs.CR

classification cs.CR
keywords topicsrealusersreproducibledatadatasetgoogleprior
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Topics API for the web is Google's privacy-enhancing alternative to replace third-party cookies. Results of prior work have led to an ongoing discussion between Google and research communities about the capability of Topics to trade off both utility and privacy. The central point of contention is largely around the realism of the datasets used in these analyses and their reproducibility; researchers using data collected on a small sample of users or generating synthetic datasets, while Google's results are inferred from a private dataset. In this paper, we complement prior research by performing a reproducible assessment of the latest version of the Topics API on the largest and publicly available dataset of real browsing histories. First, we measure how unique and stable real users' interests are over time. Then, we evaluate if Topics can be used to fingerprint the users from these real browsing traces by adapting methodologies from prior privacy studies. Finally, we call on web actors to perform and enable reproducible evaluations by releasing anonymized distributions. We find that for the 1207 real users in this dataset, the probability of being re-identified across websites is of 2%, 3%, and 4% after 1, 2, and 3 observations of their topics by advertisers, respectively. This paper shows on real data that Topics does not provide the same privacy guarantees to all users and that the information leakage worsens over time, further highlighting the need for public and reproducible evaluations of the claims made by new web proposals.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Differentially Private Synthetic Data Release for Topics API Outputs

    cs.CR 2025-06 conditional novelty 6.0 of 10

    The paper presents a differentially private methodology and a public synthetic dataset of Topics API traces that match real re-identification risk within one standard deviation on two attacks.

  2. Technical Report: The Need for a (Research) Sandstorm through the Privacy Sandbox

    cs.CR 2025-12 conditional novelty 3.0 of 10

    An independent research portal (Privacy Sandstorm) catalogs analyses and datasets on Google's Privacy Sandbox, claiming broader visibility than Google's official channels.

Pith tools