REVIEW 4 major objections 5 minor 3 cited by
The Great Data Standoff: Researchers vs. Platforms Under the Digital Services Act
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Existing TikTok data access cannot answer the key questions about the 2024 Romanian election interference, so the DSA's Article 40 route is necessary.
desk verdict Useful case-study mapping of DSA data needs for the Romanian TikTok election, but the 'unanswerable' claim leans on an under-documented audit. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the Article 40 DSA data-access pipeline, combined with the paper's availability audit. The standoff problem is the information asymmetry at the heart of the argument: researchers must request specific data without knowing what a platform collects, while platforms expect specificity and can deny that invisible data exist. The audit, summarised in Table 1, is the load-bearing instrument: each data point needed for the Romanian election case is matched against three sources, platform documentation, the TikTok Research API, and data donations, and labelled available, limited, or not available. This classification carries the argument that Article 40 is needed, because it shows which gaps are real and which are merely undocumented.
What would settle it
If a researcher with vetted access to TikTok's Research API and a donated data package from an account involved in the Romanian election could retrieve any data point the paper classifies as not available, such as algorithmic parameters, moderation flags, inferred account demographics, account type, historical engagement metrics, or suggested search terms, the central claim would lose its factual basis. A documented Article 40 request that is refused for one of these points would support it.
Extended reading notes
Core claim
The paper's central claim is that existing data access mechanisms provide insufficient means to investigate the systemic risks described in the Romanian election case study, and that access through Article 40 DSA is therefore necessary. Working from the known features of the incident, such as a coordinated network generating inauthentic engagement, undisclosed micro-influencer advertising, and livestream gifting that amplified one candidate, the paper derives concrete research tasks and the data categories they require: content, user, engagement, moderation, temporal, network, algorithmic, and geospatial data. It then classifies each data point against TikTok's privacy documentation, the Research API, and data donations obtained from the authors' own accounts. The audit shows that recommendation and moderation data are entirely inaccessible, and even within accessible categories, corresponding points such as account demographics, account type, and historical engagement metrics are missing, which the paper argues renders the research questions unanswerable through existing means.
Load-bearing premise
The audit treats the authors' own TikTok accounts, the API endpoints they selected, and public platform documentation as a complete and accurate picture of what the Research API and data donations can provide; if undocumented fields or account-dependent donation packages exist, the conclusion that the research questions are unanswerable would be too strong.
Editorial extensions
If this is right
- Researchers preparing Article 40 applications will need to justify each data request against platform documentation, as the paper demonstrates, because the audit shows that even documented data may be absent from public APIs and donations.
- Digital Services Coordinators and the European Commission must clarify data inventories under Article 6(4) of the Delegated Act, otherwise the standoff persists: researchers cannot request data they cannot see listed anywhere.
- Election-interference research on TikTok cannot proceed through compliant external means alone; the paper's classification implies that studies of recommender amplification and moderation effects will require vetted access.
- The systemic risk concept under Article 34(1)(c) DSA becomes operational only when tied to concrete data needs, as the paper does by linking electoral risk to platform manipulation and hidden advertising.
- Cross-platform research remains an open challenge because the case study is limited to TikTok, while election interference typically spans multiple platforms.
Reading between the lines
- A natural extension is to run the same documentation-to-API audit on other large platforms such as Instagram, YouTube, or X; if the same classes of data are missing, the gap is structural to the platform-researcher relationship, not specific to TikTok.
- The paper's data-need checklist could be reused as a completeness checklist for Article 40 applications, with regulators asking applicants to show which of the eight data categories each systemic-risk factor requires and where each point is available.
- A testable prediction follows from the paper's logic: if TikTok ever expands its Research API to include algorithmic parameters, moderation flags, or inferred account attributes, the unanswerable verdict on election-interference research questions should be revisited.
- The standoff might be broken not by case-by-case requests but by requiring platforms to publish machine-readable data inventories, making invisible data visible before any request is filed.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes the operationalization of Article 40 of the EU Digital Services Act (DSA), which grants vetted researchers access to very large online platforms' internal data for studying systemic risks. It argues that two obstacles impede this mechanism: the legal uncertainty of the systemic-risk concept and an information asymmetry between researchers and platforms, which the authors call the 'standoff problem.' Using the 2024 Romanian presidential election interference incident as a case study, the paper derives a set of research data needs from prior computational social science literature on platform manipulation and hidden advertising, then compares these needs against what TikTok makes available through its documentation, its Research API, and data donations. The central conclusion is that existing data-access mechanisms are insufficient to investigate the identified systemic risks and that access through Article 40 is therefore necessary. The paper is conceptual and interdisciplinary, combining legal analysis, platform governance, and computer science perspectives, and its empirical core is a detailed table (Table 1) classifying each needed data point as available, limited, or not available through the two existing access mechanisms.
Significance. If the Table 1 classification is reliable, the paper makes a useful and timely contribution: it provides a concrete, case-driven template for how researchers might frame Article 40 DSA data-access requests, and it highlights real gaps in TikTok's researcher-facing transparency tools. The combination of legal analysis with computational data-needs mapping is valuable for the platform-governance community, and the paper is honest about several limitations, including the cross-platform scope and the opacity of platform data inventories. The work does not present a formal model or machine-checked proofs, but for a conceptual contribution the clarity of the case study and the explicit mapping between research questions and data fields are strengths. The main significance, however, is conditional: the policy conclusion that Article 40 access is 'necessitated' depends on the completeness and accuracy of the availability classification in Table 1, and the methodology behind that table is currently under-documented.
major comments (4)
- [Data Availability as a Platform Governance Narrative; Table 1 (Appendix)] The paper's central claim that existing mechanisms 'provide insufficient means' to study the identified systemic risks rests entirely on the availability classification in Table 1, but the methodology behind the data-donation column is under-reported. The text states only that 'data donations ... was obtained by the authors from their own TikTok research accounts' and does not report the number of accounts, account types (personal, creator, business), regions, app versions, or the dates on which downloads were performed. TikTok's 'Download your data' exports are known to vary by account type, region, app version, and over time, so the 'No' entries for fields such as Account Gender, Account type, GPPPA status, Historical engagement metrics, Algorithmic parameters, and Suggested search terms may be artifacts of the particular accounts probed rather than properties of the data-donation mechanism as a whole. Because Article 8(6) DA requires researchers to demonstrate that data cannot be obtained through alternative existing means, this completeness claim is load-bearing; the paper should report the exact donation procedure, account characteristics, dates, and any validation against platform documentation, or soften the conclusion accordingly.
- [Data Availability as a Platform Governance Narrative; Table 1 (Appendix)] The Research API column of Table 1 is likewise presented without a systematic enumeration of the API endpoints and fields queried, the number and frequency of calls, or the dates on which the calls were made. TikTok's Research API supports an explicit 'fields' parameter that returns additional data when specified, so a 'No' entry may reflect an incomplete request rather than documented unavailability. To support the conclusion that certain data are 'not available' through the Research API, the paper should either document that every relevant endpoint and field combination was checked against current API documentation, or reclassify the affected entries as 'not examined' and adjust the central insufficiency claim accordingly.
- [Platform Manipulation and Election Interference; Political Influencer Marketing as Hidden Advertising (Summary of Data…] The data-requirements taxonomy is derived almost exclusively from studies of Twitter, Facebook, and Instagram, yet the Romanian case study itself identifies TikTok-specific mechanisms—notably livestream gifting and search-recommendation amplification—that are not mapped to any data category in the summaries. For instance, 'livestream gifting' is discussed in the case-study overview and in the preliminary evidence, but no corresponding data need (e.g., live-stream metadata, gift/coin transaction logs) appears in the 'Summary of Data Requirements' lists or in Table 1. Since the paper's conclusion is that election-related systemic risks on TikTok specifically cannot be studied with existing access mechanisms, the taxonomy should either incorporate these TikTok-specific data needs or explicitly argue that they are subsumed under the existing categories.
- [Data Availability as a Platform Governance Narrative; Article 8(6) DA] The paper's conclusion that Article 40 access is 'necessitated' relies on the premise that scraping is not a viable alternative means, but the paper dismisses scraping only by stating that it 'focus[es] on these two mechanisms that comply with the terms of service of the platform.' Article 8(6) DA requires a demonstration that the research project 'cannot be carried out with alternative existing means such as using data available through other sources,' and the paper itself notes that investigative journalists in the same case study scraped TikTok data externally. The paper should engage with the possibility that scraping or third-party tools, despite terms-of-service concerns, constitute an 'alternative existing means' for at least some of the research tasks, or explain why they are not sufficient for the scientific purposes described.
minor comments (5)
- [Platform Manipulation and Election Interference; Political Influencer Marketing as Hidden Advertising] Two section cross-references appear as empty 'Section ' placeholders in the text; these should be replaced with the appropriate section titles or numbers.
- [References] The reference list contains placeholder-like artifacts, including 'Unknown 2020', '[Title of the Press Release]' for the European Commission press release, and 'Title Unknown' for an Electrochemical Society journal reference; these must be completed or removed.
- [Conclusion] There is a typo in the final section: 'prevalant' should be 'prevalent.' Also, in the section on systemic risks under the DSA, the phrase 'content content' appears and should be corrected.
- [Figure 1] The caption describes colors from green to purple, but the text also refers to 'intermediate shades'; please ensure the figure legend is explicit enough for color-blind readers and that the caption matches the actual color mapping used in the figure.
- [Section numbering] The text refers to 'Sections 3.2 and 3.3' but the manuscript uses unnumbered section headings; update the in-text references to match the published formatting.
Circularity Check
No circular derivation; the central insufficiency claim rests on an external data-needs audit rather than on a fitted prediction or self-citation chain.
full rationale
The paper's central claim—that existing mechanisms (TikTok's Research API and data donations) provide insufficient data to investigate the Romanian election interference case—is an empirical audit, not a derivation from its own inputs. Data requirements are derived from prior literature on platform manipulation and hidden advertising (largely external to this paper, with some author contributions among many cited works), and then compared against TikTok documentation, API endpoints, and the authors' own data donations. There is no fitted parameter renamed as a prediction, no data need defined in terms of the availability classification, and no uniqueness theorem imported from the authors' prior work to force the conclusion. Self-citations such as Zannettou et al. 2023 (data donations), Bertaglia et al. 2023/2025 (sponsored-content detection), and Kaushal et al. 2024 (DSA transparency database) supply background methods and prior empirical findings, but the insufficiency conclusion rests on the Table 1 comparison, not on those citations. The main validity threat—that 'No' entries in Table 1 may reflect the authors' own account types and selected API probes rather than the full mechanisms—is a generalizability/sampling concern, not circularity: the classification is not constructed to imply the conclusion, and no specific equation or definition reduces one claim to another. Accordingly, no circular step can be exhibited under the required standard, and the paper's central argument remains self-contained against external benchmarks apart from minor reliance on the authors' own prior work for methodological framing.
Assumptions & free parameters
assumptions (3)
- domain assumption TikTok's public documentation and Research API accurately represent the data the platform collects and provides to researchers.
- domain assumption The data needs derived from prior computational social science literature are necessary and sufficient for studying platform manipulation and hidden advertising in the Romanian election case.
- domain assumption Systemic risk under DSA Article 34 is sufficiently broad to encompass election interference as studied here.
Cite this review
Pith. "Pith review of The Great Data Standoff: Researchers vs. Platforms Under the Digital Services Act." pith.science (2026). https://pith.science/paper/7GDYLLEY
@misc{pith2026250501122,
author = {Pith},
title = {Pith review of: The Great Data Standoff: Researchers vs. Platforms Under the Digital Services Act},
year = {2026},
howpublished = {\url{https://pith.science/paper/7GDYLLEY}},
note = {Machine review of arXiv:2505.01122}
}
read the original abstract
To facilitate accountability and transparency, the Digital Services Act (DSA) sets up a process through which Very Large Online Platforms (VLOPs) need to grant vetted researchers access to their internal data (Article 40(4)). Operationalising such access is challenging for at least two reasons. First, data access is only available for research on systemic risks affecting European citizens, a concept with high levels of legal uncertainty. Second, data access suffers from an inherent standoff problem. Researchers need to request specific data but are not in a position to know all internal data processed by VLOPs, who, in turn, expect data specificity for potential access. In light of these limitations, data access under the DSA remains a mystery. To contribute to the discussion of how Article 40 can be interpreted and applied, we provide a concrete illustration of what data access can look like in a real-world systemic risk case study. We focus on the 2024 Romanian presidential election interference incident, the first event of its kind to trigger systemic risk investigations by the European Commission. During the elections, one candidate is said to have benefited from TikTok algorithmic amplification through a complex dis- and misinformation campaign. By analysing this incident, we can comprehend election-related systemic risk to explore practical research tasks and compare necessary data with available TikTok data. In particular, we make two contributions: (i) we combine insights from law, computer science and platform governance to shed light on the complexities of studying systemic risks in the context of election interference, focusing on two relevant factors: platform manipulation and hidden advertising; and (ii) we provide practical insights into various categories of available data for the study of TikTok, based on platform documentation, data donations and the Research API.
Figures
Forward citations
Cited by 3 Pith papers
-
Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times
Under the DSA, researchers still face slow, opaque, and often unsuccessful platform data access applications, with rejection or silence common.
-
Algorithmic Authority and the Clinical Standard of Care
Proposes a dialectical standard of care that integrates AI and physicians as a single accountable unit, using Lessig's framework and an analogy between algorithmic and human errors.
-
Improving Regulatory Oversight in Online Content Moderation
The paper proposes a cross-checking process between DSA Transparency Database records and Transparency Reports, plus a verification process against platform-held data, to detect inconsistencies in self-reported moderation.
Reference graph
Works this paper leans on
-
[1]
For most authors... (a) Would answering this research question advance sci- ence without violating social contracts, such as violat- ing privacy norms, perpetuating unfair profiling, exac- erbating the socio-economic divide, or implying disre- spect to societies or cultures? Yes. This is a conceptual paper that analyses the DSA’s Article 40 data-access me...
-
[2]
InProceedings of the 19th ACM Asia Conference on Computer and Communications Security, 1784–1800
SoK: False Information, Bots and Malicious Cam- paigns: Demystifying Elements of Social Media Manipula- tions. InProceedings of the 19th ACM Asia Conference on Computer and Communications Security, 1784–1800. Allcott, H.; and Gentzkow, M. 2017. Social media and fake news in the 2016 election.Journal of economic perspectives, 31(2): 211–236. Badawy, A.; Fe...
work page 2017
-
[3]
(a) Did you state the full set of assumptions of all theoret- ical results? NA
Additionally, if you are including theoretical proofs... (a) Did you state the full set of assumptions of all theoret- ical results? NA. (b) Did you include complete proofs of all theoretical re- sults? NA
-
[4]
Additionally, if you ran machine learning experiments... (a) Did you include the code, data, and instructions needed to reproduce the main experimental results (ei- ther in the supplemental material or as a URL)? NA. (b) Did you specify all the training details (e.g., data splits, hyperparameters, how they were chosen)? NA. (c) Did you report error bars (...
-
[5]
(a) If your work uses existing assets, did you cite the cre- ators? NA
Additionally, if you are using existing assets (e.g., code, data, models) or curating/releasing new assets,without compromising anonymity... (a) If your work uses existing assets, did you cite the cre- ators? NA. (b) Did you mention the license of the assets? NA. (c) Did you include any new assets in the supplemental material or as a URL? NA. (d) Did you ...
-
[6]
(a) Did you include the full text of instructions given to participants and screenshots? NA
Additionally, if you used crowdsourcing or conducted research with human subjects,without compromising anonymity... (a) Did you include the full text of instructions given to participants and screenshots? NA. (b) Did you describe any potential participant risks, with mentions of Institutional Review Board (IRB) ap- provals? NA. (c) Did you include the est...
-
[7]
(a) Did you clearly state the assumptions underlying all theoretical results? NA
Additionally, if your study involves hypotheses testing... (a) Did you clearly state the assumptions underlying all theoretical results? NA. (b) Have you provided justifications for all theoretical re- sults? NA. (c) Did you discuss competing hypotheses or theories that might challenge or complement your theoretical re- sults? NA. (d) Have you considered ...
-
[2020]
A Survey on Computational Propaganda Detection
Characterizing social media manipulation in the 2020 US presidential election.First Monday. Golovchenko, Y .; Buntain, C.; Eady, G.; Brown, M. A.; and Tucker, J. A. 2020. Cross-platform state propaganda: Rus- sian trolls on twitter and YouTube during the 2016 US Presi- dential Election.The International Journal of Press/Politics, 25(3): 357–389. Grinberg,...
work page Pith review arXiv 2020
Show all 11 references
-
[2021]
In2021 IEEE European symposium on security and privacy (EuroS&P), 403–422
Ephemeral astroturfing attacks: The case of fake twit- ter trends. In2021 IEEE European symposium on security and privacy (EuroS&P), 403–422. IEEE. Ershov, D.; and Mitchell, M. 2020. The Effects of Influ- encer Advertising Disclosure Regulations: Evidence From Instagram. InPro...
2020 arXiv
-
[2024]
Akhtar, M
Accessed: 2025-01-22. Akhtar, M. M.; Masood, R.; Ikram, M.; and Kanhere, S. S
2025
-
[2025]
Bertaglia, T.; Huber, S.; Goanta, C.; Spanakis, G.; and Iamnitchi, A
Influencer self-disclosure practices on Instagram: A multi-country longitudinal study.Online Social Networks and Media, 45: 100298. Bertaglia, T.; Huber, S.; Goanta, C.; Spanakis, G.; and Iamnitchi, A. 2023. Closing the Loop: Testing ChatGPT to Generate Model Explanations to I...
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.