{"id":"c3c304fe-1ab4-42db-9cd9-6869f6355374","arxiv_id":"2505.09877","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Under the DSA, researchers still face slow, opaque, and often unsuccessful platform data access applications, with rejection or silence common.","lead":"Researchers who study social media report that the European Union's new transparency rules have not fixed data access: applications are slow, opaque, and often rejected. A survey of 180 researchers and interviews with 19 suggest the promise of the Digital Services Act is not yet reaching the people who study platforms.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-selection and missing response-rate data make the survey's prevalence claims about DSA data-access inadequacy fragile; the caveat in Section 5 explicitly disclaims representativeness, so the central claim outruns the evidence.","rationale":"The paper's qualitative findings are credible and align with previous API audits, so this is not a rejection. The load-bearing weakness is the quantitative generalization: the central claim uses prevalence language, but the survey is based on a purposive, self-selected sample with no reported response rate, and the authors explicitly disclaim representativeness. The reader identified the same sampling limitation; the stress-test adds a concrete falsification strategy, a non-response follow-up, that would either confirm or mitigate the selection-bias concern. Because the evidence supports a conditional verdict rather than a firm rejection or acceptance, the reader's CONDITIONAL verdict remains appropriate, and no verdict change is recommended.","tokens_in":22052,"tokens_out":6555,"duration_ms":72447,"concrete_test":"Run a non-response follow-up: send a short version of the survey with the same key items on awareness, application, and outcomes to a random sample of non-respondents drawn from the same seven professional organizations' mailing lists, and compare their response distributions with the original 180. If non-respondents report materially lower barriers, the original sample is selection-biased and prevalence claims should be downgraded; if distributions are similar, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that DSA-era data access programs are 'far from adequate' requires evidence about how widespread and severe the barriers are. The survey uses 180 volunteer respondents recruited through professional organizations and community lists, with no invitation denominator, no response rate, and no full instrument; the 19 interviewees are also volunteers, with two recruited by snowball referral. Section 5 concedes that 'the experiences of the researchers involved in this study are not representative of any research field or region.' If researchers who had negative experiences were more likely to respond, Tables 1 and 2 will overstate unawareness, problematic applications, and denials. The qualitative interviews independently establish that serious barriers exist, so the paper does not collapse; but the prevalence component of the headline claim is not supported. Table 2 also omits a 'pending' column, so the raw counts cannot directly verify the text's claim that most applications were pending.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This mixed-methods paper examines researchers' experiences with platform data access programs during the early implementation of the EU Digital Services Act (DSA), a period the authors label the 'post-post-API age.' The authors fielded a survey (n=180) through professional organizations and conducted 19 semi-structured interviews with researchers (mostly academic, EU- and US-based) from October 2024 to February 2025. They report that researchers face barriers at three stages: applying (low awareness, complex forms, IRB requirements), obtaining access (opaque rejections, long delays, exclusion of non-academic researchers), and using the data (poor API quality, restrictive caps, inaccurate data, especially for TikTok). They conclude that DSA-mandated data access programs are 'far from adequate' and offer recommendations for platforms, researchers, and policymakers. The paper includes a survey summary (Tables 1-2), a flowchart of the access process, and a participant table.","tokens_in":22131,"tokens_out":11432,"duration_ms":103730,"significance":"If the findings hold, the paper is a timely and valuable empirical contribution to CSCW/HCI and internet governance debates. The qualitative arm is a clear strength: 19 interviews with pseudonyms, direct quotations, and transparent coding procedures (open coding, codebook, axial coding) produce a rich account of the practical shortcomings of DSA-era data access. The cross-platform scope (X, TikTok, Meta, YouTube, etc.) is broader than most prior audits, and the recommendations for platforms, researchers, and policymakers are concrete. The paper also frankly acknowledges its purposive sampling limits in Section 5. However, the quantitative component is weaker: the survey's non-probability sample and incomplete reporting (no response rate, no pending column in Table 2, omission of the promised Reddit data) mean the paper's headline prevalence claims should be read as describing the volunteer respondents, not the population of researchers. With appropriate rescoping, the paper's central qualitative conclusion is well supported.","major_comments":[{"comment":"The survey was distributed through seven professional organizations plus four research communities, but the paper reports no invitation denominator, no response rate, and no full instrument, and Section 5 concedes that the sample is 'not representative of any research field or region.' The quantitative prevalence claims in Section 4.1 (e.g., 'many researchers did not apply' because unaware or found applications problematic) therefore describe only the 180 volunteer respondents, yet the Discussion opens by asserting generally that 'current data access programs are far from being adequate.' Please rescope the survey-based claims to the sample (or provide response rates and a formal analysis), while keeping the qualitative conclusions, which are independently supported.","section":"Section 3.1 and Section 5"},{"comment":"The application-outcome table reports Applied, Denied, and Accepted counts but has no explicit 'Pending' column, so the text's statement that 'the majority of applications were still pending' can only be verified by subtraction (e.g., 64 of 116 total applications are unaccounted for if all remaining rows imply pending). The subsequent claim that researchers waited 'at least a month' with 'some' waiting longer is not supported by any reported waiting-time statistics. Please add a pending column and report waiting-time frequencies or ranges.","section":"Section 4.1, Table 2"},{"comment":"The methods state that the survey 'included an additional platform, Reddit,' and Reddit is discussed in the interviews (e.g., Kay and Thirteen in Section 4.2.2), but Reddit appears in neither Table 1 nor Table 2. If Reddit survey data were collected, they should be reported; if not, the text should be corrected to avoid a factual inconsistency.","section":"Section 4.1 and Tables 1-2"},{"comment":"The survey mixes platforms that actually maintain DSA Article 40 researcher access programs with platforms that do not (e.g., Alibaba, Booking.com, Google Shopping) and includes Reddit, which the authors note is not a VLOP. As a result, the 'Unaware' and 'Not interested' counts for platforms without researcher access programs do not speak to the adequacy of DSA-mandated access, and the aggregate table can overstate unawareness across the DSA ecosystem. Please either restrict the primary analysis to platforms with identifiable Article 40 programs or disaggregate results by program type.","section":"Section 4.1 and Background Section 2.3"}],"minor_comments":[{"comment":"'emperical' is a typo for 'empirical' (appears twice in the Conclusion); 'effectivly' in Section 4.2.2 should be 'effectively.'","section":"Section 6"},{"comment":"The survey instrument is not provided; including it as an appendix or supplementary material would improve replicability.","section":"Section 3.1"},{"comment":"The text says 'nearly all of the interviewees who have used the TikTok API' reported problems; please state how many of the 19 interviewees had actually used the TikTok API.","section":"Section 4.2.3"},{"comment":"The label 'Crowdtangle' should be 'CrowdTangle' for consistency, and the note 'Includes Facebook and Instagram' should clarify whether it covers both CrowdTangle and Meta Content Library rows.","section":"Table 1"},{"comment":"The statement that 'many cases appear to fall under the DSA's scope according to both our analysis and participants' perspectives' is an interpretive legal judgment; consider citing the specific Article 40 criteria applied or framing it explicitly as participants' perceptions.","section":"Section 5.1"},{"comment":"The 'Post-Post-API' era is dated 2023-, which aligns with the DSA Article 40 entry into force, but the text in Section 2.3 describes the regulatory developments without specifying a year; a small clarification would help readers.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and addresses a timely issue. The main revisions concern the quantitative reporting; if the authors can add the pending data, correct the Reddit omission, and temper the generalization in the Discussion, the paper could be a strong contribution. I would not reject: the qualitative evidence alone supports the core conclusion that DSA data-access programs are currently inadequate in practice. The authors' explicit non-representativeness caveat is to their credit, but it makes the current overstatement in the Discussion more puzzling."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth engaging. Its real contribution is the first cross-platform, mixed-methods account of researcher experiences with DSA Article 40 data access programs. The interviews with 19 researchers are the heart: named pseudonyms, direct quotes, and a transparent coding process. They document barriers at every stage—unawareness, onerous and opaque applications, unexplained denials, and APIs that are clunky or return inconsistent data after access is granted. That evidence alone supports the central claim that current programs fall far short of the DSA's promise.\n\nThe survey component is the weakest part. We get raw counts by platform but no response rate, no invitation denominator, no full instrument, and no statistical treatment. The sample is purposive and self-selected, and the authors explicitly say in Section 5 that the findings are not representative of any field or region. The stress-test worry about prevalence is only partly fair: the paper does use sample-bound phrases like \"many\" and \"majority,\" but the Discussion's core assertion is about the programs' inadequacy, which the interviews establish independently. A fixable reporting gap: Table 2 has no \"pending\" column even though the text says most applications were pending.\n\nCitation practice looks fine—prior API audits and legal analyses are acknowledged, and the authors' earlier work is background, not load-bearing. No code or data is shipped, which is understandable for interviews but a data-availability statement would still help.\n\nVerdict: a solid, useful empirical snapshot. I'd send it to peer review and ask for revisions: share the survey instrument, report response rates and denominators, add the pending column, and soften the quantitative generalizations to match the sample. The qualitative strand does not need to change.","headline":"A timely mixed-methods snapshot of how researchers actually fare under DSA data access programs; the interview evidence is strong and the survey is thin but honestly caveated.","tokens_in":22723,"tokens_out":1645,"would_cite":true,"duration_ms":18261,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that DSA-mandated data-access programs are not providing researchers with adequate platform data in practice.","keywords":["post-post-API age","Digital Services Act","DSA Article 40","platform data access","researcher access programs","social media API","data transparency","mixed-method study"],"falsifier":"A complete audit of DSA Article 40 applications submitted to the major very large online platforms, including application counts, wait times, approval rates, denial reasons, and a ground-truth check of returned API data against the platforms' own public interfaces. If approvals are prompt and widespread, denials are explained, and API data matches observable public content, the paper's central claim of inadequacy would be contradicted.","tokens_in":21787,"feed_emoji":"🔍","tokens_out":5776,"duration_ms":54906,"temperature":0.7,"pith_summary":"The paper argues that the data-access programs platforms built to comply with the European Union's Digital Services Act are not, in practice, giving researchers the data they need. It supports this with a survey of 180 researchers and in-depth interviews with 19, documenting failures at each stage: researchers do not know the programs exist, applications are cumbersome and slow, approvals are often denied or met with silence, and the APIs that do work return incomplete or inconsistent data. If the paper is right, the regulatory promise of transparency has not yet materialized, and platform-based social media research remains fragile.","feed_headline":"DSA data access is falling short, researcher survey finds","feed_subtitle":"180 researchers and 19 interviews report rejections, silence, and clunky APIs from regulated platforms.","key_machinery":"The paper's core mechanism is the 'post-post-API age' frame paired with a mixed-method study. It divides platform data access into four eras—pre-API, voluntary-API, post-API, and post-post-API—and treats the Digital Services Act's Article 40 researcher-access mandates as the defining feature of the newest era. Methodologically, it combines a survey of 180 researchers with 19 semi-structured interviews and organizes the results around a flowchart of four barriers: availability and awareness, application, access, and usability. This four-stage pipeline is what carries the argument that a researcher must clear every hurdle to conduct research.","core_discovery":"The central claim is that DSA-mandated public data access is falling well short of the law's intent. Across very large online platforms and search engines, researchers encounter four compounding barriers—unawareness or disinterest, problematic application processes, long or absent responses, and inadequate API usability—that push many to abandon studies or turn to scraping and third-party tools. The paper states plainly that despite platforms' efforts to meet data transparency requirements, practices vary greatly and current data access programs are far from adequate to facilitate research on digital platforms.","pith_inferences":["If the documented pattern persists, social media research will increasingly skew toward well-funded academic groups in the US and EU, because the application and usability barriers function as filters on who can do the work.","A natural next step would be a longitudinal public dashboard of per-platform application outcomes, allowing regulators and the community to track whether data access improves or worsens over time.","The 'independence by permission' critique implies that even a smoother application process would not resolve the underlying problem: platforms still decide which research questions are permissible, so external appeal mechanisms matter more than procedural tweaks.","If official APIs remain unusable, user-centric methods such as data donation and tracking may become the main route for independent research, trading breadth for consent and control."],"forward_implications":["If current DSA access programs remain inadequate, the transparency regulation will not generate the independent platform research it was designed to enable.","Researchers without university affiliation, EU-based status, or funding face disproportionate barriers, so institutional, regional, and financial inequities in data access widen.","The difficulty of official access pushes researchers toward web scraping and paid third-party vendors, which carry legal ambiguity and inconsistent data quality.","Opaque denial decisions, especially from X and TikTok, leave researchers unable to contest outcomes and erode trust in platform-provided access.","Regulatory clarification, such as the European Commission's planned Delegated Regulation, is needed to define who qualifies and what data must be shared."],"supporting_citations":[{"why":"Documents how platform APIs are designed for business objectives rather than academic inquiry, establishing the structural control that underlies access barriers.","marker":"[13]"},{"why":"Defines the 'post-API age' and the loss of compliant data collection, which the paper extends into its era framework.","marker":"[23]"},{"why":"Supplies the 'independence by permission' concept describing researchers' dependence on corporate approval for platform data.","marker":"[42]"},{"why":"Provides the prior audit of TikTok's research API showing data inaccuracies, which the paper's interviews corroborate.","marker":"[35]"},{"why":"Documents the legal and operational uncertainties of DSA data access, used to explain researchers' confusion about eligibility and available data.","marker":"[24]"},{"why":"Argues that DSA Article 40 covers non-academic researchers, forming the basis for the paper's claim that platforms exclude them unfairly.","marker":"[20]"},{"why":"Describes the post-API turn to unsanctioned scraping and platforms' fight against it, contextualizing the alternative approaches researchers adopt.","marker":"[12]"},{"why":"Early report on TikTok's researcher API problems, referenced as prior evidence that the API is hard to use.","marker":"[10]"}],"fun_headline_variants":["Post-API promise unmet: DSA data access still scarce","DSA data access: more barriers, fewer insights","Researchers hit walls in DSA data access quest","DSA data access promises, but researchers get silence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The findings rest on a purposive, self-selected sample of 180 survey respondents and 19 interviewees; if frustrated researchers were more likely to respond than those who gained access easily, the barriers would appear more common than they are.","fun_headline_variants_meta":{"raw":{"variants":["Post-API promise unmet: DSA data access still scarce","DSA data access: more barriers, fewer insights","Researchers hit walls in DSA data access quest","DSA data access promises, but researchers get silence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000827,"raw_usage":{"total_tokens":3569,"prompt_tokens":856,"completion_tokens":2713,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":2649}},"tokens_in":472,"tokens_out":2713,"duration_ms":19512,"temperature":1.0,"reasoning_tokens":2649,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:21:33.163844+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A complete audit of DSA Article 40 applications submitted to the major very large online platforms, including application counts, wait times, approval rates, denial reasons, and a ground-truth check of returned API data against the platforms' own public interfaces. If approvals are prompt and widespread, denials are explained, and API data matches observable public content, the paper's central claim of inadequacy would be contradicted.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents how platform APIs are designed for business objectives rather than academic inquiry, establishing the structural control that underlies access barriers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 'independence by permission' concept describing researchers' dependence on corporate approval for platform data."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the prior audit of TikTok's research API showing data inaccuracies, which the paper's interviews corroborate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues that DSA Article 40 covers non-academic researchers, forming the basis for the paper's claim that platforms exclude them unfairly."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Early report on TikTok's researcher API problems, referenced as prior evidence that the API is hard to use."}],"review_version":1}