REVIEW 4 major objections 6 minor 2 cited by
Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that DSA-mandated data-access programs are not providing researchers with adequate platform data in practice.
desk verdict A timely mixed-methods snapshot of how researchers actually fare under DSA data access programs; the interview evidence is strong and the survey is thin but honestly caveated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's core mechanism is the 'post-post-API age' frame paired with a mixed-method study. It divides platform data access into four eras—pre-API, voluntary-API, post-API, and post-post-API—and treats the Digital Services Act's Article 40 researcher-access mandates as the defining feature of the newest era. Methodologically, it combines a survey of 180 researchers with 19 semi-structured interviews and organizes the results around a flowchart of four barriers: availability and awareness, application, access, and usability. This four-stage pipeline is what carries the argument that a researcher must clear every hurdle to conduct research.
What would settle it
A complete audit of DSA Article 40 applications submitted to the major very large online platforms, including application counts, wait times, approval rates, denial reasons, and a ground-truth check of returned API data against the platforms' own public interfaces. If approvals are prompt and widespread, denials are explained, and API data matches observable public content, the paper's central claim of inadequacy would be contradicted.
Extended reading notes
Core claim
The central claim is that DSA-mandated public data access is falling well short of the law's intent. Across very large online platforms and search engines, researchers encounter four compounding barriers—unawareness or disinterest, problematic application processes, long or absent responses, and inadequate API usability—that push many to abandon studies or turn to scraping and third-party tools. The paper states plainly that despite platforms' efforts to meet data transparency requirements, practices vary greatly and current data access programs are far from adequate to facilitate research on digital platforms.
Load-bearing premise
The findings rest on a purposive, self-selected sample of 180 survey respondents and 19 interviewees; if frustrated researchers were more likely to respond than those who gained access easily, the barriers would appear more common than they are.
Editorial extensions
If this is right
- If current DSA access programs remain inadequate, the transparency regulation will not generate the independent platform research it was designed to enable.
- Researchers without university affiliation, EU-based status, or funding face disproportionate barriers, so institutional, regional, and financial inequities in data access widen.
- The difficulty of official access pushes researchers toward web scraping and paid third-party vendors, which carry legal ambiguity and inconsistent data quality.
- Opaque denial decisions, especially from X and TikTok, leave researchers unable to contest outcomes and erode trust in platform-provided access.
- Regulatory clarification, such as the European Commission's planned Delegated Regulation, is needed to define who qualifies and what data must be shared.
Reading between the lines
- If the documented pattern persists, social media research will increasingly skew toward well-funded academic groups in the US and EU, because the application and usability barriers function as filters on who can do the work.
- A natural next step would be a longitudinal public dashboard of per-platform application outcomes, allowing regulators and the community to track whether data access improves or worsens over time.
- The 'independence by permission' critique implies that even a smoother application process would not resolve the underlying problem: platforms still decide which research questions are permissible, so external appeal mechanisms matter more than procedural tweaks.
- If official APIs remain unusable, user-centric methods such as data donation and tracking may become the main route for independent research, trading breadth for consent and control.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This mixed-methods paper examines researchers' experiences with platform data access programs during the early implementation of the EU Digital Services Act (DSA), a period the authors label the 'post-post-API age.' The authors fielded a survey (n=180) through professional organizations and conducted 19 semi-structured interviews with researchers (mostly academic, EU- and US-based) from October 2024 to February 2025. They report that researchers face barriers at three stages: applying (low awareness, complex forms, IRB requirements), obtaining access (opaque rejections, long delays, exclusion of non-academic researchers), and using the data (poor API quality, restrictive caps, inaccurate data, especially for TikTok). They conclude that DSA-mandated data access programs are 'far from adequate' and offer recommendations for platforms, researchers, and policymakers. The paper includes a survey summary (Tables 1-2), a flowchart of the access process, and a participant table.
Significance. If the findings hold, the paper is a timely and valuable empirical contribution to CSCW/HCI and internet governance debates. The qualitative arm is a clear strength: 19 interviews with pseudonyms, direct quotations, and transparent coding procedures (open coding, codebook, axial coding) produce a rich account of the practical shortcomings of DSA-era data access. The cross-platform scope (X, TikTok, Meta, YouTube, etc.) is broader than most prior audits, and the recommendations for platforms, researchers, and policymakers are concrete. The paper also frankly acknowledges its purposive sampling limits in Section 5. However, the quantitative component is weaker: the survey's non-probability sample and incomplete reporting (no response rate, no pending column in Table 2, omission of the promised Reddit data) mean the paper's headline prevalence claims should be read as describing the volunteer respondents, not the population of researchers. With appropriate rescoping, the paper's central qualitative conclusion is well supported.
major comments (4)
- [Section 3.1 and Section 5] The survey was distributed through seven professional organizations plus four research communities, but the paper reports no invitation denominator, no response rate, and no full instrument, and Section 5 concedes that the sample is 'not representative of any research field or region.' The quantitative prevalence claims in Section 4.1 (e.g., 'many researchers did not apply' because unaware or found applications problematic) therefore describe only the 180 volunteer respondents, yet the Discussion opens by asserting generally that 'current data access programs are far from being adequate.' Please rescope the survey-based claims to the sample (or provide response rates and a formal analysis), while keeping the qualitative conclusions, which are independently supported.
- [Section 4.1, Table 2] The application-outcome table reports Applied, Denied, and Accepted counts but has no explicit 'Pending' column, so the text's statement that 'the majority of applications were still pending' can only be verified by subtraction (e.g., 64 of 116 total applications are unaccounted for if all remaining rows imply pending). The subsequent claim that researchers waited 'at least a month' with 'some' waiting longer is not supported by any reported waiting-time statistics. Please add a pending column and report waiting-time frequencies or ranges.
- [Section 4.1 and Tables 1-2] The methods state that the survey 'included an additional platform, Reddit,' and Reddit is discussed in the interviews (e.g., Kay and Thirteen in Section 4.2.2), but Reddit appears in neither Table 1 nor Table 2. If Reddit survey data were collected, they should be reported; if not, the text should be corrected to avoid a factual inconsistency.
- [Section 4.1 and Background Section 2.3] The survey mixes platforms that actually maintain DSA Article 40 researcher access programs with platforms that do not (e.g., Alibaba, Booking.com, Google Shopping) and includes Reddit, which the authors note is not a VLOP. As a result, the 'Unaware' and 'Not interested' counts for platforms without researcher access programs do not speak to the adequacy of DSA-mandated access, and the aggregate table can overstate unawareness across the DSA ecosystem. Please either restrict the primary analysis to platforms with identifiable Article 40 programs or disaggregate results by program type.
minor comments (6)
- [Section 6] 'emperical' is a typo for 'empirical' (appears twice in the Conclusion); 'effectivly' in Section 4.2.2 should be 'effectively.'
- [Section 3.1] The survey instrument is not provided; including it as an appendix or supplementary material would improve replicability.
- [Section 4.2.3] The text says 'nearly all of the interviewees who have used the TikTok API' reported problems; please state how many of the 19 interviewees had actually used the TikTok API.
- [Table 1] The label 'Crowdtangle' should be 'CrowdTangle' for consistency, and the note 'Includes Facebook and Instagram' should clarify whether it covers both CrowdTangle and Meta Content Library rows.
- [Section 5.1] The statement that 'many cases appear to fall under the DSA's scope according to both our analysis and participants' perspectives' is an interpretive legal judgment; consider citing the specific Article 40 criteria applied or framing it explicitly as participants' perceptions.
- [Figure 1] The 'Post-Post-API' era is dated 2023-, which aligns with the DSA Article 40 entry into force, but the text in Section 2.3 describes the regulatory developments without specifying a year; a small clarification would help readers.
Circularity Check
No circularity: conclusions rest on independent self-reported data; self-citations are background only.
full rationale
This paper is an empirical mixed-methods study (180 survey responses, 19 interviews), not a derivation chain. The central claim—that DSA-era data access programs are "far from being adequate to facilitate research on digital platforms" (Section 5)—is a summary of participants' self-reported experiences and the raw counts in Tables 1 and 2. No fitted parameter is renamed as a prediction, no quantity is defined in terms of another, and no uniqueness theorem, ansatz, or first-principles result is imported via self-citation. The only self-references are background or corroborative: [10] is cited to align with interviewees' complaints about TikTok's API ("The platform most frequently mentioned was TikTok, aligning with previous reports [10]"), [11] is cited for web-scraping legal/ethical context, and [22] is described as an example infrastructure. Removing any of these citations would not change the paper's conclusion, so none is load-bearing. The paper itself flags the main validity limitation in Section 5: "both our qualitative and quantitative studies relied on purposive sampling, and the experiences of the researchers involved in this study are not representative of any research field or region." That caveat limits generalizability but is not a form of circularity. Table 2 omits a 'pending' column even though the text states that "the majority of applications were still pending"; this is a reporting/verifiability issue, not a circular-reasoning issue. No circular steps are exhibited.
Assumptions & free parameters
assumptions (3)
- domain assumption Self-reports from researchers accurately reflect platform application processes and outcomes.
- domain assumption DSA Article 40.12 is the correct legal benchmark for public data access.
- domain assumption The sample can reveal the existence of barriers even if it cannot estimate their prevalence.
Cite this review
Pith. "Pith review of Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times." pith.science (2026). https://pith.science/paper/2LHG6LEF
@misc{pith2026250509877,
author = {Pith},
title = {Pith review of: Post-Post-API Age: Studying Digital Platforms in Scant Data Access Times},
year = {2026},
howpublished = {\url{https://pith.science/paper/2LHG6LEF}},
note = {Machine review of arXiv:2505.09877}
}
read the original abstract
Over the past decade, data provided by digital platforms has informed substantial research in HCI to understand online human interaction and communication. Following the closure of major social media APIs that previously provided free access to large-scale data (the "post-API age"), emerging data access programs required by the European Union's Digital Services Act (DSA) have sparked optimism about increased platform transparency and renewed opportunities for comprehensive research on digital platforms, leading to the "post-post-API age." However, it remains unclear whether platforms provide adequate data access in practice. To assess how platforms make data available under the DSA, we conducted a comprehensive survey followed by in-depth interviews with 19 researchers to understand their experiences with data access in this new era. Our findings reveal significant challenges in accessing social media data, with researchers facing multiple barriers including complex API application processes, difficulties obtaining credentials, and limited API usability. These challenges have exacerbated existing institutional, regional, and financial inequities in data access. Based on these insights, we provide actionable recommendations for platforms, researchers, and policymakers to foster more equitable and effective data access, while encouraging broader dialogue within the CSCW community around interdisciplinary and multi-stakeholder solutions.
Figures
Forward citations
Cited by 2 Pith papers
-
Auditing Meta and TikTok Research API Data Access under Article 40(12) of the Digital Services Act
TikTok and Meta research APIs expose only about 75% and 50% of user-visible posts, respectively, and strip most contextual metadata, making independent auditing of systemic risks structurally biased.
-
The Big Ban Theory: A Pre- and Post-Intervention Dataset of Online Content Moderation Actions
A new public dataset aligns pre- and post-intervention user activity for 25 Reddit/Voat moderation actions, enabling comparative moderation research.
Reference graph
Works this paper leans on
-
[1]
Davey Alba. 2019. Ahead of 2020, Facebook falls short on plan to share data on disinformation. https://www.nytimes.com/2019/09/29/technology/facebook-disinformation.html [Accessed 2025-05-01]
work page 2019
-
[2]
Davey Alba. 2021. Facebook sent flawed data to misinformation researchers. https://www.nytimes.com/live/2020/2020- election-misinformation-distortions#facebook-sent-flawed-data-to-misinformation-researchers [Accessed 2025-05-01]
work page 2021
-
[3]
Jennifer Allen, Markus Mobius, David M Rothschild, and Duncan J Watts. 2021. Research note: Examining potential bias in large-scale censored data.Harvard Kennedy School Misinformation Review(2021)
work page 2021
-
[4]
Adriana Alvarado Garcia, Tianling Yang, and Milagros Miceli. 2025. What Knowledge Do We Produce from Social Media Data and How?Proceedings of the ACM on Human-Computer Interaction9, 1 (2025), 1–45
work page 2025
-
[5]
Theo Araujo, Jef Ausloos, Wouter van Atteveldt, Felicia Loecherbach, Judith Moeller, Jakob Ohme, Damian Trilling, Bob van de Velde, Claes De Vreese, and Kasper Welbers. 2022. OSD2F: An open-source data donation framework. Computational Communication Research4, 2 (2022), 372–387
work page 2022
-
[6]
Ayesha Bhimdiwala, Krishna Akhil Kumar Adavi, and Ahmer Arif. 2024. Fighting for Their Voice: Understanding Indian Muslim Women’s Responses to Networked Harassment.Proceedings of the ACM on Human-Computer Interaction 8, CSCW1 (2024), 1–24
work page 2024
-
[7]
Elizabeth Blakey. 2024. The Day Data Transparency Died: How Twitter/X Cut Off Access for Social Research.Contexts 23, 2 (2024), 30–35
work page 2024
-
[8]
Shannon Bond. 2023. Elon Musk sues disinformation researchers, claiming they are driving away advertis- ers. https://www.npr.org/2023/08/01/1191318468/elon-musk-sues-disinformation-researchers-claiming-they-are- driving-away-adverti [Accessed 2025-05-01]
work page 2023
Show all 44 references
-
[9]
Johannes Breuer, Zoltán Kmetty, Mario Haim, and Sebastian Stier. 2023. User-centric approaches for collecting Facebook data in the ‘post-API age’: Experiences from two studies and recommendations for future research.Information, Communication & Society26, 14 (2023), 2649–2668
2023
-
[10]
Megan A. Brown. 2023. The Problem with TikTok’s New Researcher API is Not TikTok. https://www.techpolicy.press/the-problem-with-tiktoks-new-researcher-api-is-not-tiktok [Accessed 2025-05- 01]
2023
-
[11]
Brown, Andrew Gruen, Gabe Maldoff, Solomon Messing, Zeve Sanderson, and Michael Zimmer
Megan A. Brown, Andrew Gruen, Gabe Maldoff, Solomon Messing, Zeve Sanderson, and Michael Zimmer. 2024. Web Scraping for Research: Legal, Ethical, Institutional, and Scientific Considerations. arXiv:2410.23432 [cs.CY] https://arxiv.org/abs/2410.23432
2024 arXiv
-
[12]
Axel Bruns. 2021. After the ‘APIcalypse’: Social media platforms and their fight against critical scholarly research. Disinformation and data lockdown on social platforms(2021), 14–36
2021
-
[13]
Taina Bucher. 2013. Objects of intense feeling: The case of the Twitter API.Computational Culture3 (2013)
2013
-
[14]
Thijs C Carrière, Laura Boeschoten, Bella Struminskaya, Heleen L Janssen, Niek C de Schipper, and Theo Araujo. 2025. Best practices for studies using digital data donation.Quality & Quantity59, 1 (2025), 389–412. doi:10.1007/s11135- 024-01983-x
2025 doi
-
[15]
Peter Chapman. 2024. Laying the Foundation for Independent Platform Data Access in the EU. https://kgi.georgetown.edu/research-and-commentary/independent-platform-data-access-eu [Accessed 2025-05-01]
2024
-
[16]
Emily Chen, Kristina Lerman, Emilio Ferrara, et al . 2020. Tracking social media discourse about the COVID-19 pandemic: Development of a public coronavirus twitter data set.JMIR public health and surveillance6, 2 (2020), e19273
2020
-
[17]
Clara Christner, Aleksandra Urman, Silke Adam, and Michaela Maier. 2022. Automated Tracking Approaches for Studying Online Media Use: A Critical Review and Recommendations.Communication Methods and Measures16, 2 (2022), 79–95. doi:10.1080/19312458.2021.1907841
2022
-
[18]
Justine Clama. 2023. Twitter just closed the book on academic research. https://www.theverge.com/2023/5/31/23739084/twitter-elon-musk-api-policy-chilling-academic-research [Ac- cessed 2025-05-01]
2023
-
[19]
Francesco Corso, Francesco Pierri, and Gianmarco De Francisci Morales. 2024. What we can learn from TikTok through its Research API. InCompanion Publication of the 16th ACM Web Science Conference(Stuttgart, Germany)(Websci Companion ’24). Association for Computing Machinery, N...
2024
-
[20]
Mateus Correia de Carvalho. 2024. Researcher Access to Platform Data and the DSA: One Step Forward, Three Steps Back. https://www.techpolicy.press/researcher-access-to-platform-data-and-the-dsa-one-step-forward-three-steps- back [Accessed 2025-04-28]. J. ACM, Vol. 37, No. 4, A...
2024
-
[21]
European Commission. 2023. Delegated Regulation on data access provided for in the Digital Services Act. https://ec.europa.eu/info/law/better-regulation/have-your-say/initiatives/13817-Delegated-Regulation-on-data- access-provided-for-in-the-Digital-Services-Act_en [Accessed 2...
2023
-
[22]
Alvaro Feal, Jeffrey Gleason, Pranav Goel, Jason Radford, Kai-Cheng Yang, John Basl, Michelle Meyer, David Choffnes, Christo Wilson, and David Lazer. 2024. Introduction to National Internet Observatory. InProceedings of ICWSM Data Challenge Workshop(Buffalo, NY, USA). AAAI
2024
-
[23]
Deen Freelon. 2018. Computational research in the post-API age.Political Communication35, 4 (2018), 665–668
2018
-
[24]
Catalina Goanta, Savvas Zannettou, Rishabh Kaushal, Jacob van de Kerkhof, Thales Bertaglia, Taylor Annabell, Haoyang Gui, Gerasimos Spanakis, and Adriana Iamnitchi. 2025. The Great Data Standoff: Researchers vs. Platforms Under the Digital Services Act. arXiv:2505.01122 [cs.CY...
2025 arXiv
-
[25]
Rae, Oriol Vinyals, and Laurent Sifre
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osin...
2022 arXiv
-
[26]
2020.# HashtagActivism: Networks of race and gender justice
Sarah J Jackson, Moya Bailey, and Brooke Foucault Welles. 2020.# HashtagActivism: Networks of race and gender justice. MIT Press
2020
-
[27]
Abhinav Reddy Karra, Ranjan Jaiswal, and Sanorita Dey. 2023. Fishing for Validation: Understanding Promises and Challenges of a Private Social Media Group for COVID-19 Long-Hauler Patients.Proceedings of the ACM on Human-Computer Interaction7, CSCW1 (2023), 1–34
2023
-
[28]
Gary King and Nathaniel Persily. 2020. A new model for industry–academic partnerships.PS: Political Science & Politics 53, 4 (2020), 703–709
2020
-
[29]
David Lazer, Alex Pentland, Lada Adamic, Sinan Aral, Albert-László Barabási, Devon Brewer, Nicholas Christakis, Noshir Contractor, James Fowler, Myron Gutmann, et al. 2009. Computational social science.Science323, 5915 (2009), 721–723
2009
-
[30]
Meta Platforms, Inc. n.d.. Meta Content Library and API. https://developers.facebook.com/docs/content-library-and- api/content-library [Accessed 2025-04-28]
2025
-
[31]
Ryan Murtfeldt, Naomi Alterman, Ihsan Kahveci, and Jevin D. West. 2024. RIP Twitter API: A eulogy to its vast research contributions. arXiv:2404.07340 [cs.CY] https://arxiv.org/abs/2404.07340
2024
-
[32]
Reeves, and Thomas N
Jakob Ohme, Theo Araujo, Laura Boeschoten, Deen Freelon, Nilam Ram, Byron B. Reeves, and Thomas N. Robinson and. 2024. Digital Trace Data Collection for Social Media Effects Research: APIs, Data Donation, and (Screen) Tracking. Communication Methods and Measures18, 2 (2024), 1...
2024
-
[33]
Barbara Ortutay. 2024. Meta kills off misinformation tracking tool CrowdTangle despite pleas from re- searchers, journalists. https://apnews.com/article/meta-crowdtangle-research-misinformation-shutdown-facebook- 977ece074b99adddb4887bf719f2112a [Accessed 2025-05-01]
2024
-
[34]
Jessica A Pater, Amanda Coupe, Fayika Farhat Nova, Rachel Pfafman, Jeanne Carroll, Abigal Brouwer, Camden Bohn, Jason Li, Noah Todd, Fen Lei Chang, et al . 2023. Social Media is Not a Health Proxy: Differences Between Social Media and Electronic Health Record Reports of Post-C...
2023
-
[35]
George DH Pearson, Nathan A Silver, Jessica Y Robinson, Mona Azadi, Barbara A Schillo, and Jennifer M Kreslake. 2025. Beyond the margin of error: A systematic and replicable audit of the TikTok research API.Information, Communication & Society28, 3 (2025), 452–470
2025
-
[36]
2020.Social media and democracy: The state of the field, prospects for reform
Nathaniel Persily, Joshua A Tucker, and Joshua Aaron Tucker. 2020.Social media and democracy: The state of the field, prospects for reform. Cambridge University Press
2020
-
[37]
Stiene Praet, David Martens, and Peter Van Aelst. 2021. Patterns of democracy? Social network analysis of parliamentary Twitter networks in 12 countries.Online Social Networks and Media24 (2021), 100154
2021
-
[38]
Stephen Prochaska, Kayla Duskin, Zarine Kharazian, Carly Minow, Stephanie Blucker, Sylvie Venuto, Jevin D West, and Kate Starbird. 2023. Mobilizing manufactured reality: How participatory disinformation shaped deep stories to catalyze action during the 2020 US presidential ele...
2023
-
[39]
Nicholas Proferes, Naiyan Jones, Sarah Gilbert, Casey Fiesler, and Michael Zimmer. 2021. Studying reddit: A systematic overview of disciplines, approaches, methods, and ethics.Social Media+ Society7, 2 (2021), 20563051211019004
2021
-
[40]
Kate Starbird, Ahmer Arif, and Tom Wilson. 2019. Disinformation as collaborative work: Surfacing the participatory nature of strategic information operations.Proceedings of the ACM on human-computer interaction3, CSCW (2019), 1–26
2019
-
[41]
u/PeerRevue. [n. d.]. R/reddit4researchers. https://www.reddit.com/r/reddit4researchers/?rdt=35953
-
[42]
Michael W Wagner. 2023. Independence by permission.Science381, 6656 (2023), 388–391. J. ACM, Vol. 37, No. 4, Article 111. Publication date: May 2025. 111:24 Kayo Mimizuka, Megan A Brown, Kai-Cheng Yang, and Josephine Lukito Table 3. Participant information. Pseudonym Pronoun A...
2023
-
[43]
Yihe Wang and Kathryn E Ringland. 2023. Weaving Autistic Voices on TikTok: Utilizing Co-Hashtag Networks for Netnography. InCompanion Publication of the 2023 Conference on Computer Supported Cooperative Work and Social Computing. 254–258
2023
-
[44]
Angela Xiao Wu and Harsh Taneja. 2021. Platform enclosure of human behavior and its measurement: Using behavioral trace data against platform episteme.New Media & Society23, 9 (2021), 2650–2667. A Participant Information In Table 3, we provide detailed information about the pa...
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.