REVIEW 5 major objections 5 minor 29 references
Auditing Meta and TikTok Research API Data Access under Article 40(12) of the Digital Services Act
T0 review · 5 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Researchers auditing Meta and TikTok through official APIs see only a fraction of the public posts users actually encounter, because platform-imposed filters exclude up to half of posts and strip most contextual metadata.
desk verdict A genuinely useful first cut at quantifying DSA API data loss, but the headline percentages are less stable than the paper suggests. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the audit design: the authors define the 'public information environment' (PIE) as the full set of content and metadata transmitted to users' devices during normal platform use, then capture that baseline with controlled sockpuppet accounts that record complete HTTP payloads of recommended posts. Against this baseline, the same posts are queried through the official Research APIs (Meta Content Library, TikTok Research API). Differences are classified into three named filters that carry the argument: scope narrowing (posts excluded by account thresholds, format, moderation, or ingestion delay), metadata stripping (contextual fields removed), and operational restr
What would settle it
Replicate the audit with non-political, general-interest sockpuppet accounts during a non-election period on the same platforms: if the Research APIs then retrieve close to 100 percent of user-visible posts with high metadata coverage, the claim of structural bias would be refuted. A second check: have a platform release a complete log of all public posts recommended to a panel of consenting human users during a fixed window and compare it against the API's retrievable subset; if the API contains all of those posts, the 'up to ~50% loss' finding fails.
Extended reading notes
Core claim
The paper's central claim is that the Meta Content Library and TikTok Research API—the primary instruments for DSA-mandated researcher access—return a structurally incomplete and biased representation of the platforms' public information environment (PIE), defined as all content and metadata delivered to users' devices during normal use. Benchmarking full feeds captured from two sockpuppet accounts against the same posts queried through the Research APIs, the authors show that researchers can retrieve roughly 75 percent (TikTok) and 50 percent (Instagram) of user-visible posts, and only 17 percent and 42 percent of the metadata parameters transmitted with those posts. This data loss is not r
Load-bearing premise
The comparison rests on the assumption that the two politically oriented sockpuppet accounts' feeds faithfully represent the platform's public information environment; if platforms treat automated accounts differently or election-period content is unrepresentative, the measured loss percentages may not generalize to typical user experience.
Editorial extensions
If this is right
- Research on elections, misinformation, or moderation that draws only on Meta and TikTok Research APIs will systematically undercount removed, ephemeral, and rapidly amplified content, because the surviving dataset is biased toward content that stayed online long enough to be archived.
- Findings will not be stable over time: content deletion, moderation, and mandatory data-refresh requirements mean the same API query about the same period yields different results on different days, undermining replication.
- Platforms' self-reported compliance metrics and any regulator analysis built on their Research APIs inherit the survivorship bias documented here, so observed 'absence of evidence' of systemic risks cannot be read as evidence of absence.
- Legal definitions of 'publicly accessible data' under the DSA must be anchored in the user-visible public information environment rather than in platform-administered thresholds, or accountability becomes procedural.
- The three identified filters (scope, metadata, operational) collectively mean that even a vetted researcher with full API credentials cannot reconstruct what an ordinary user encountered, making a core premise of Article 40(12) study unfulfillable in practice.
Reading between the lines
- We would infer, beyond the paper's claim, that the same sockpuppet-audit protocol could be extended to other very large platforms (e.g., YouTube, X, LinkedIn) to test whether comparable data-loss structures appear under their Article 40(12) implementations.
- The metadata fields observed in browser traffic but absent from APIs (e.g., moderation decision flags, AI-detection markers) suggest a testable path: repeated sockpuppet runs over time could reveal when platforms change their amplification and moderation interventions—information the APIs currently render invisible.
- A direct test of generalizability: run the same audit with non-political, general-interest accounts outside election windows; high API coverage there would localize the bias to electoral contexts, while persistent losses would support the paper's structural conclusion.
- The paper's reform proposal to let vetted researchers scrape public data could be evaluated by simulating a 'compliant API' that returns the full PIE with metadata, and then measuring how much of the observed data loss actually disappears; this would separate the filters that are removable by policy from those inherent to platform architecture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper audits whether the TikTok Research API and the Meta Content Library can reconstruct the public information environment (PIE) that users actually see. Using two sockpuppet accounts—one trained on U.S. election content and one on German election content—the authors captured the TikTok For You feed and the Instagram Explore feed with their own SOAP tool, then attempted to retrieve the same posts through the official Research APIs. They report that only about 75% (TikTok) and 50% (Instagram) of user-visible posts are accessible through the APIs, and that for accessible posts only 17% (TikTok) and 42% (Instagram) of metadata parameters are retained. The authors attribute this to three overlapping mechanisms: scope narrowing (account-size thresholds, ephemeral/live formats, moderation and deletion, ingestion delays), metadata stripping (missing labels, outdated engagement data), and operational restrictions (rate limits, data-enclave rules). They conclude that the APIs are structurally incomplete and biased and propose amendments to Article 40(12), including a broader definition of publicly accessible data, full contextual metadata, operationally effective access, and permission for researchers to scrape public data. The paper is open about many limitations but does not test how much its quantitative results generalize.
Significance. The paper addresses a timely and policy-relevant question: whether official research APIs are fit for the DSA's systemic-risk auditing purpose. Its strengths are the direct user-centric audit design, the use of open-source tooling, the explicit mechanism taxonomy (scope narrowing, metadata stripping, operational restrictions), and a largely self-contained comparison with no fitted parameters or circular definitions. If the reported magnitudes are robust, the paper provides valuable empirical evidence for ongoing regulatory enforcement debates. However, the quantitative core currently rests on exactly two sockpuppet feeds, both election-tuned, and on an unspecified API-matching protocol. The headline 'up to ~50%' also appears to depend on a historical account-size threshold rather than the threshold in effect during the Instagram data collection. The contribution is promising, but the central quantitative claims need additional robustness analysis before they can support the paper's broad structural-bias conclusion.
major comments (5)
- [§3.1, §5.1] The central quantitative result is computed from one TikTok For You feed and one Instagram Explore feed, both from accounts deliberately trained on election content (§3.1). Section 5.1 acknowledges that magnitudes may vary by context and that platforms might treat sockpuppets differently, but it does not test this. No second account with different training, no non-political feed, and no per-account dispersion is reported. The rebuttal that sockpuppets only measure a differential between user-facing and researcher-facing data assumes that differential is invariant to account type and content domain; that invariance is precisely what needs evidence. Please provide at least a small robustness panel (e.g., additional accounts and a non-election feed) or explicitly restrict the conclusions to election-related, algorithmically surfaced content.
- [§3.2, §4 first paragraph] The paper does not specify how individual PIE posts were matched to Research API records. It is never stated which API endpoints were queried (by post/video ID, by account, by keyword), which fields served as join keys, how absent/deleted/unavailable responses were coded, at what dates API queries were run relative to feed collection, or how API errors were handled. Without this protocol, the 75%/50% access rates and the 17%/42% metadata ratios are not independently reproducible, which is a serious gap for an audit paper. Please provide a precise matching and classification procedure, including query templates, join keys, and handling of non-retrieval.
- [§4.1.1, Table 1] Table 1 shows that the Instagram excluded-post share varies from 49.35% to 11.45% to 2.47% depending on which Meta follower threshold is applied. The Instagram feed was collected from 08 Jan to 25 Feb 2025, after Meta's Version 5.0 had already lowered the threshold to 1,000 followers. The abstract's 'up to approximately 50 percent' therefore reflects a historical 25,000-follower policy epoch, not the access regime in effect during the Instagram data collection. This is a load-bearing distinction: the headline overstates current scope loss and conflates a discretionary policy choice with a stable property of the API. Please report the actually applicable threshold as primary and clearly label the historical counterfactual.
- [§4.2.1, Table 3] The metadata retention ratios are computed by counting parameters in HTTP responses and comparing them with documented API fields. This treats all parameters as equally meaningful and conflates raw payload size with analytical information loss. Table 3 itself labels many fields as 'potential relation' and infers semantics from naming conventions. An exploratory inventory is fine, but the claim that the APIs strip 'essential contextual metadata' needs a more substantive analysis: which omitted fields are needed for concrete systemic-risk research questions, and which are redundant technical fields? Without such grounding, the 17%/42% figures overstate the strength of the metadata-stripping evidence.
- [§4, §5, Conclusion] The paper repeatedly concludes that research APIs are 'structurally biased,' but the only direct distributional comparison is for deleted/unavailable TikTok posts, where the authors show that high-reach posts are affected (§4.1.3). The authors do not systematically compare characteristics (topic, engagement, account size, moderation outcomes, posting time) between matched accessible and inaccessible posts for either platform. To support 'structural bias,' please provide such a comparison, or soften the conclusion to 'incomplete and potentially biased in ways that are not yet quantified across covariates.' This is important because the paper's central contribution is not merely that data are missing, but that missingness is systematically correlated with risk-relevant content.
minor comments (5)
- [General] The running header contains a placeholder author name ('Trovato and Tobin, et al.') that should be replaced with the actual authors.
- [§3.1] Typo: 'sockuppet-swiping' should be 'sockpuppet-swiping.' Also, the sentence 'This choice is intentional, as election-related content exemplifies...' is repeated or rephrased in §5.1; consider consolidating.
- [Fig. 1] The figure legend uses n_HTTP=236, n_MCL-API=100, n_MCL-UI=14 but does not define how nested objects or repeated fields are counted. Please clarify the counting scheme and add axis labels or a caption describing the data source (one Instagram post).
- [Table 2] The percentages in the rows do not always sum to exactly 100% (e.g., 17 Feb: 82.27+6.89+10.84=100.00; 19 Feb: 81.51+7.53+10.97=100.01; 04 Mar: 76.66+10.59+12.76=100.01). Also, define operationally what distinguishes 'Deleted' from 'Unavailable' (e.g., archived, private, removed by platform).
- [References] Reference [15] is cited as 'Knight Georgetown Institute' but the official name is the Knight-Georgetown Institute (KGI). Reference [14] appears to duplicate the author name 'Gabor Halasz.' Please correct the reference metadata.
Circularity Check
No significant circularity: the audit compares independently captured user-visible feeds with API responses; the headline quantities are measurements, not derived predictions.
full rationale
The paper's central chain is empirical rather than definitional: SOAP captures the full HTTP payloads delivered to two sockpuppet feeds (the PIE baseline), the researchers then query the official Meta and TikTok Research APIs for the same posts, and the reported percentages (approximately 75%/50% post accessibility, 17%/42% metadata retention) are direct counts from that comparison. No parameter is fitted to the target result and then relabeled as a prediction; the account-threshold analysis in Table 1 applies documented platform thresholds to the observed follower distribution and is presented as a scenario analysis, not as a fitted finding. The PIE definition is explicitly stated and applied consistently: content and metadata delivered to a user's device constitute the measured baseline, and the metadata-stripping figures are an operationalization of that definition rather than a conclusion smuggled in from an external assumption. The main self-citations ([3,4]) introduce the SOAP measurement tool, but the tool is open source and reproducible, and the paper does not rely on a self-cited uniqueness theorem or ansatz to force its conclusions. The acknowledged limitation regarding sockpuppet representativeness is a generalization/validity concern, not circularity: the authors explicitly defend the design as measuring a differential between user-conveyed and researcher-conveyed PIE, and this defense does not reduce the measured differential to the paper's own assumptions. Overall, the derivation chain is self-contained and externally checkable, so no circular step meets the required evidentiary standard.
Assumptions & free parameters
assumptions (3)
- domain assumption The sockpuppet accounts' feeds are treated as representative of the platform's public information environment (PIE).
- domain assumption Every HTTP response parameter transmitted to a user's browser counts as 'publicly accessible data'.
- domain assumption A post not retrievable through the Research API is absent due to platform-imposed filters, not due to a querying artifact.
Cite this review
Pith. "Pith review of Auditing Meta and TikTok Research API Data Access under Article 40(12) of the Digital Services Act." pith.science (2026). https://pith.science/paper/D62SKMCO
@misc{pith2026260112390,
author = {Pith},
title = {Pith review of: Auditing Meta and TikTok Research API Data Access under Article 40(12) of the Digital Services Act},
year = {2026},
howpublished = {\url{https://pith.science/paper/D62SKMCO}},
note = {Machine review of arXiv:2601.12390}
}
read the original abstract
Article 40(12) of the Digital Services Act (DSA) requires Very Large Online Platforms (VLOPs) to provide vetted researchers with access to publicly accessible data. While prior work has identified shortcomings of platform-provided data access mechanisms, existing research has not quantitatively assessed data quality and completeness in Research APIs across platforms, nor systematically mapped how current access provisions fall short. This paper presents a systematic audit of research access modalities by comparing data obtained through platform Research APIs with data collected about the same platforms' user-visible public information environment (PIE). Focusing on two major platform APIs, the TikTok Research API and the Meta Content Library, we reconstruct full information feeds for two controlled sockpuppet accounts during two election periods and benchmark these against the data retrievable for the same posts through the corresponding Research APIs. Our findings show systematic data loss through three classes of platform-imposed mechanisms: scope narrowing, metadata stripping, and operational restrictions. Together, these mechanisms implement overlapping filters that exclude large portions of the platform PIE (up to approximately 50 percent), strip essential contextual metadata (up to approximately 83 percent), and impose severe technical constraints for researchers (down to approximately 1000 requests per day). Viewed through a data quality lens, these filters primarily undermine completeness, resulting in a structurally biased representation of platform activity. We conclude that, in their current form, the Meta and TikTok Research APIs fall short of supporting meaningful, independent auditing of systemic risks as envisioned under the DSA.
Figures
Reference graph
Works this paper leans on
-
[1]
Ortega, Manuel Álvarez-Mon, and Rocío M
Irene Alfonso-Fuertes, María Álvarez-Mon, Raquel Sánchez del Hoyo, Miguel A. Ortega, Manuel Álvarez-Mon, and Rocío M. Molina-Ruiz. 2023. Time Spent on Instagram and Body Image, Self-esteem, and Physical Comparison Among Young Adults in Spain: Observational Study.JMIR Formative Research7 (2023), e42207. doi:10.2196/42207
-
[2]
Tolulope Balogun. 2025. Preserving Digital Footprints: Strategies for Safeguarding Ephemeral Online Data.International Journal of Knowledge Content Development & Technology15, 1 (Apr. 2025), 19–32. https://ijkcdt.journals.publicknowledgeproject.org/index.php/ijkcdt/article/view/1075
2025
-
[3]
Luka Bekavac, Kimberly Garcia, Jannis Strecker, Simon Mayer, and Aurelia Tamo-Larrieux. 2024. From Walls to Windows: Creating Transparency to Understand Filter Bubbles in Social Media. InNORMalize 2024: The Second Workshop on the Normative Design and Evaluation of Recommender Systems, co-located with the ACM Conference on Recommender Systems 2024 (RecSys ...
2024
-
[4]
Luka Bekavac, Jannis Strecker-Bischoff, Kimberly Garcia, Simon Mayer, and Aurelia Tamò-Larrieux. 2026. Scrutinizing Systemic Risks in Personalized Recommender Systems Through Sock-Puppet Auditing of VLOPs.ACM Transactions on Recommender Systems(2026). To appear
2026
-
[5]
Mark Bernstein. 2022. The Web At War: Hypertext, Social Media, and Totalitarianism: Hypertext, Social Media, and Totalitarianism. InProceedings of the 33rd ACM Conference on Hypertext and Social Media(Barcelona, Spain)(HT ’22). Association for Computing Machinery, New York, NY, USA, 256–258. doi:10.1145/3511095.3536365
arXiv 2022
-
[6]
Axel Bruns. 2019. After the ‘APIcalypse’: social media platforms and their fight against critical scholarly research.Information, Communication & Society22, 11 (2019), 1544–1566. doi:10.1080/1369118X.2019.1637447 arXiv:https://doi.org/10.1080/1369118X.2019.1637447
arXiv 2019
-
[7]
Kate Conger and Ryan Mac. 2024. Musk’s Trump talk on X: After glitchy start, a Two-Hour ramble.The New York Times(Aug. 2024). https: //www.nytimes.com/2024/08/13/technology/elon-musk-x-donald-trump.html
2024
-
[8]
Aron Culotta. 2010. Towards detecting influenza epidemics by analyzing Twitter messages. InProceedings of the First Workshop on Social Media Analytics(Washington D.C., District of Columbia)(SOMA ’10). Association for Computing Machinery, New York, NY, USA, 115–122. doi:10.1145/ 1964858.1964874
arXiv 2010
Show all 29 references
-
[9]
Carlos Entrena-Serrano, Martin Degeling, Salvatore Romano, and Raziye Buse Çetin. 2025. TikTok’s Research API: Problems Without Explanations. arXiv:2506.09746 [cs.CY] https://arxiv.org/abs/2506.09746
2025 arXiv
-
[10]
European Commission. 2024. Status Report: Mechanisms for Researcher Access to Online Platform Data. https://digital-strategy.ec.europa.eu/en/ library/status-report-mechanisms-researcher-access-online-platform-data
2024
-
[11]
European Union. 2022. Regulation (EU) 2022/2065 of the European Parliament and of the Council on a Single Market for Digital Services (Digital Services Act). Official Journal of the EU L277/1
2022
-
[12]
Mozilla Foundation. 2024. https://www.mozillafoundation.org/en/blog/new-research-tech-platforms-data-access-initiatives-vary-widely/
2024
-
[13]
Mozilla Foundation. 2025. https://www.mozillafoundation.org/en/what-we-do/mobilize/fair-terms-report/
2025
-
[14]
Gabor Halasz and Gabor Halasz. 2025. Gespräch ohne Widerspruch zwischen Musk und Weidel.tagesschau.de(Dec. 2025). https://www.tagesschau. de/inland/bundestagswahl/parteien/weidel-musk-100.html
2025
-
[15]
Knight Georgetown Institute. 2025. https://kgi.georgetown.edu/research-and-commentary/better-access/
2025
-
[16]
Julian Jaursch, Jakob Ohme, and Ulrike Klinger. 2024. Enabling Research with Publicly Accessible Platform Data: Early DSA Compliance Issues and Suggestions for Improvement. https://www.weizenbaum-library.de/items/39a490c3-f3c1-42a7-a530-318b17e9de49
2024
-
[17]
Publicly Accessible
Daphne Keller. 2025.The Stakes of “Publicly Accessible”: Researchers’ Rights to Data under the DSA. https://verfassungsblog.de/dsa-fine-x-research- data/ VerfBlog
2025
-
[18]
Anja Lambrecht and Catherine Tucker. 2019. Algorithmic Bias? An Empirical Study of Apparent Gender-Based Discrimination in the Display of STEM Career Ads.Management Science65 (04 2019). doi:10.1287/mnsc.2018.3093
2019
-
[19]
Kayo Mimizuka, Megan A Brown, Kai-Cheng Yang, and Josephine Lukito. 2025. Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times. arXiv:2505.09877 [cs.HC] https://arxiv.org/abs/2505.09877
2025 arXiv
-
[20]
Myers, Chenguang Zhu, and Jure Leskovec
Seth A. Myers, Chenguang Zhu, and Jure Leskovec. 2012. Information diffusion and external influence in networks. InProceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining(Beijing, China)(KDD ’12). Association for Computing Machinery, ...
2012
-
[21]
Pearson, Nathan A
George D.H. Pearson, Nathan A. Silver, Jessica Y. Robinson, Mona Azadi, Barbara A. Schillo, and Jennifer M. Kreslake. 2025. Beyond the margin of error: a systematic and replicable audit of the TikTok research API.Information, Communication & Society28, 3 (2025), 452–470. doi:1...
2025
-
[22]
Jean-Christophe Plantin, Carl Lagoze, Paul N Edwards, and Christian Sandvig. 2018. Infrastructure studies meet platform studies in the age of Google and Facebook.New Media & Society20, 1 (2018), 293–310. doi:10.1177/1461444816661553 arXiv:https://doi.org/10.1177/1461444816661553
2018 doi
-
[23]
Santiago Sordo Ruz, Martin Degeling, and Kathy Meßmer. 2023. The Research API falls woefully short. https://tiktok-audit.com/blog/2023/the- TikTok-research-API-falls-woefully-short/ Last accessed December 13, 2024. Manuscript submitted to ACM 16 Trovato and Tobin, et al
2023
-
[24]
Zélia Raposo Santos, Christy M K Cheung, Pedro Simões Coelho, and Paulo Rita. 2022. Consumer engagement in social media brand communities: A literature review.International Journal of Information Management63 (2022), 102457. doi:10.1016/j.ijinfomgt.2021.102457
2022
-
[25]
2025.Data Access for Researchers under the Digital Services Act: From Policy to Practice
LK Seiling, Iglesias Keller Clara, Jakob Ohme, De Vreese Claes, and Ulrike Klinger. 2025.Data Access for Researchers under the Digital Services Act: From Policy to Practice. doi:10.34669/wi.wpp/14
2025 doi
-
[26]
2018.The Platform Society
José van Dijck, Thomas Poell, and Martijn de Waal. 2018.The Platform Society. Oxford University Press. doi:10.1093/oso/9780190889760.001.0001
2018
-
[27]
Gummadi, Elissa M
Savvas Zannettou, Olivia-Nemes Nemeth, Oshrat Ayalon, Angelica Goetzen, Krishna P. Gummadi, Elissa M. Redmiles, and Franziska Roesner
-
[28]
Schäfer, and Thiago M
Jing Zeng, Mike S. Schäfer, and Thiago M. Oliveira. 2022. Conspiracy theories in digital environments: Moving the research field forward.Convergence: The International Journal of Research into New Media Technologies28, 4 (Aug. 2022), 929–939. doi:10.1177/13548565221117474 Epub...
2022 doi
-
[2024]
arXiv:2301.04945 [cs.SI] https: //arxiv.org/abs/2301.04945
Analyzing User Engagement with TikTok’s Short Format Video Recommendations using Data Donations. arXiv:2301.04945 [cs.SI] https: //arxiv.org/abs/2301.04945
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.