{"id":"ed658f03-b3f4-4fb5-ab56-0ec4acec2f52","arxiv_id":"2506.09746","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 2025 audit of about 260,000 donated TikTok links measured that around 12% of still-public videos cannot be retrieved through TikTok's Research API, with ads, official TikTok content, and some entire accounts systematically missing.","lead":"An audit of TikTok's Research API found that it fails to return metadata for roughly one in eight public videos from donated user data, including the platform's own posts and advertisements. This matters because the API is TikTok's main legal route for independent researchers under Europe's Digital Services Act, and the authors argue that data scraping is needed to keep it honest.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 12.46% headline statistic is not reproducible from the paper's own counts: the text reports 18,961 IDs still not retrievable, but the summary implies ~32,310 public-but-unavailable videos; the 18%/36%/46% breakdown also conflicts with the '83%' phrase.","rationale":"The reader's verdict was CONDITIONAL, citing both arithmetic problems and the unstated scraping ground truth. I agree with CONDITIONAL, but the single most load-bearing concern is the internal accounting of the headline 12.46% figure. This concern does not depend on external assumptions about the scraper or sample representativeness: even if every scraper classification were perfect, the paper's own reported counts cannot be reconciled. The 12,373 recovered total, the 18,961 still-not-retrievable count, and the summary's 18%/36%/46% breakdown cannot all describe the same 70,239-ID set, and the report does not explain which number is the basis for the 12.46% claim. This is not an accusation of fabrication; it is likely a reporting or extrapolation gap, but it is precisely the kind of gap that prevents a reader from verifying the central quantitative contribution. The qualitative conclusion that the API is incomplete is supported by the concrete examples (official TikTok videos, ads, and accounts such as Taylor Swift's), so a full rejection would be too harsh. Requiring the authors to release the ID-level audit trail and reconcile the counts is an appropriate condition for acceptance. I therefore do not change the reader's CONDITIONAL verdict.","tokens_in":9588,"tokens_out":10164,"duration_ms":114971,"concrete_test":"Publish the full ID-level audit trail for the 70,239 initially missing videos: for each ID, record the initial batch result, the individual re-query outcome (pending, retrieved, not-retrieved), the scraper classification (public, not-public, unavailable, unknown), and the assigned category (Canada, ad, account-level, minor, none). Then recompute the summary flow: total recovered, total confirmed public-but-not-retrievable after all completed rechecks, and the headline rate equal to confirmed_public_unavailable divided by 260,000. If the confirmed count is 18,961, the rate is about 7.3%, not 12.46%; if the rate is 12.46%, the report's '18,961' figure or its category percentages must be corrected. This single accounting reconciliation settles whether the central statistic is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's main quantitative finding is the 'one in eight' (12.46%) failure rate, but the Experiment 1 narrative and the Summary of findings give irreconcilable numbers. The experiment starts with 70,239 initially non-retrieved IDs; after batch and individual re-querying, the text reports a total of 12,373 recovered and '18961 IDs are still not retrievable' (p. 8). The summary then states that only 18% of the 70,239 were retrieved, 36% were deleted/private, and 46% were public but unavailable via the API, yielding the 12.46% figure. Yet 18,961 is 27% of 70,239, not 46%; 46% corresponds to roughly 32,310 videos, a gap of about 13,350 from the directly reported still-not-retrievable count. Moreover, 18%+36%+46%=100%, while the text says '83 percent that remained unavailable', so the base for these percentages is ambiguous. The final 12.46% matches 46% of 70,239 divided by 260,000, not the 18,961 confirmed IDs. Thus the headline rate appears to be a scraper-derived projection rather than the observed re-query result, or the '18,961' figure is only a subset and the report never says so. The self-identified note that the test was interrupted and the authors were 'continuing to recheck the posts left' (p. 8) confirms the re-querying was incomplete, but no extrapolation rule is given. This internal inconsistency is load-bearing because the paper's central contribution is precisely this statistic; if the confirmed count is 18,961, the rate would be about 7.3%, not 12.46%.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports an empirical audit of TikTok's Research API. Using a list of approximately 260,000 TikTok URLs obtained through data donations, the authors attempted to retrieve metadata via the official API. They report that 70,239 posts initially returned no metadata; after batch and individual re-querying they recovered 12,373 posts, leaving 18,961 IDs still non-retrievable. They scraped the platform to classify non-retrievable posts, finding that about 36% were deleted/private and that a large share were publicly available but not returned by the API. They highlight several categories of excluded content (TikTok company videos, advertisements, and accounts such as Canadian creators), report a second scraping-based experiment, and maintain a public dashboard monitoring 10 unavailable videos. The central claim is that the API is unreliable for data-donation research, with roughly one in eight donated videos missing metadata with no explanation.","tokens_in":160,"tokens_out":6859,"duration_ms":138371,"significance":"The central claim is policy-relevant and the paper provides a needed field test of a regulated data-access mechanism; if the numbers were correct, the result would support the position that the TikTok Research API is not yet DSA-compliant for data-donation workflows. The paper contributes a large-scale empirical dataset, an external scraping benchmark, a reproducible failure-detection approach, and a publicly available monitoring dashboard. However, the main headline statistic is not internally consistent with the raw counts provided, and the paper does not release its data or scraper code; the quantitative basis for the policy conclusion therefore needs substantial clarification and repair.","major_comments":[{"comment":"The paper's headline statistic is not reproducible from its own counts. In Experiment 1, after batch re-querying, 'Approximately 32,000 TikTok videos remained for testing'; after individual re-querying, 12,373 additional posts were retrieved and '18961 IDs are still not retrievable' (p. 8). Yet the Summary of findings states that only 18% of the 70,239 were retrieved and that 46% (approximately 32,310) of videos were public but unavailable via the API (p. 13). The directly observed remaining count is 18,961, which is 27% of 70,239, not 46%. The 46% figure appears to derive from an extrapolation or from the scraped classification of the initial non-retrieved set, but no rule for extrapolation is given, and the relation to the 18,961 re-query result is never explained. If the observed count is used, the failure rate is 18,961/260,000, approximately 7.3%, not the claimed 12.46%. Because the 'one in eight' figure is the central quantitative finding, this inconsistency is load-bearing and must be resolved.","section":"Experiment 1 / Summary of findings (pp. 7–8, 13)"},{"comment":"The percentage breakdown in the summary is internally inconsistent. The text reports that 18% of videos were retrieved and then refers to 'the 83 percent that remained unavailable' (18% + 83% = 101%). It then attributes 36% of the unavailable videos to deletion/private status and 46% to public-but-unavailable, and adds 'the remaining 21% of videos' for which the reason is unknown; 36% + 46% + 21% = 103% if all are percentages of the same base. If instead the 21% is a sub-share of the 46% category (e.g., Canada 12%, ads 13%, unknown 21%), that should be stated explicitly and Figure 6 should define its denominator unambiguously. As written, the figure cannot be read without guessing which shares are shares of what.","section":"Summary of findings (p. 13)"},{"comment":"The conclusion asserts that 'almost 10,000 advertisements' are not accessible through the API, but no table, formula, or classification step in the paper leads to this number. The only ad-related quantitative detail in the experiments is the qualitative discussion of Figures 3–5 and the absence of a count in the 'Advertisements' section. If this number comes from the 46% category or from a separate analysis of the dashboard, that source must be specified, and the sample size used for the estimate should be reported.","section":"Conclusion (p. 17)"},{"comment":"The generalizability of the headline rate to data-donation research depends on how the 260,000-URL sample was constructed, but the paper provides no donor-recruitment details. The authors do not state who the donors were, how their data was collected, or whether the donated list over-represents particular creators, regions, or video types. Similarly, the scraper used to classify 'public' versus 'not public' is not described: no rules for handling age-restricted, geo-blocked, or regionally unavailable videos, and no validation of the scraper against the API. Without these details, the 12.46% figure cannot be interpreted as a general property of the API rather than a property of this particular donation sample.","section":"Experiment 1 (pp. 6–7)"}],"minor_comments":[{"comment":"The number '18961' should be written as '18,961' for readability.","section":"p. 8"},{"comment":"'12,46%' uses a decimal comma; the rest of the paper uses dots, so this should be '12.46%'.","section":"p. 13"},{"comment":"The 62.7% figure for publicly available posts in Experiment 1 and the 46% public-but-unavailable figure in the summary are never reconciled; the paper should state whether these are measured at different stages of the re-query process.","section":"pp. 7 and 13"},{"comment":"The caption says 'Videos available ... within the first 100 videos' but the axis describes the number of creators; clarify the unit being plotted and how the percentages were computed.","section":"Figure 7 (p. 14)"},{"comment":"The reference to 'Daikeler et al. (2024)' is not accompanied by a full bibliographic entry; please add it to the references.","section":"Known Limitations (p. 5)"},{"comment":"The dashboard methodology should be documented in the paper or on the dashboard itself, including query timing, exact API endpoints used, error handling, and how the 10 monitored videos were selected.","section":"Monitoring APIs (pp. 15–16)"}],"recommendation":"major_revision","confidential_remarks":"The paper is written as a policy report and contains several arithmetic inconsistencies that are quickly fixable; the underlying measurements seem valuable. The authors should be asked to provide exact per-category counts and a clear description of the extrapolation from re-query results to the 12.46% headline, and to correct the percentage arithmetic before further review."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The subject is important and the qualitative findings are plausible, but you cannot trust the headline number as written. The paper's central claim is that TikTok's Research API fails to return metadata for one in eight donated videos (12.46%). That statistic does not reconcile with the counts reported in the text, and the paper never explains the gap.\n\nWhat is genuinely new here: the data-donation test with roughly 260,000 URLs is a realistic audit scenario, and the exclusion categories are concrete and useful—TikTok's own official videos, advertisements, Canadian accounts, and a set of about 1% of creators whose videos are entirely absent. The batch-corruption observation (a single non-public ID causing still-public IDs in the same batch to be dropped) is new and practically important for anyone using batch queries. The dashboard is a reasonable ongoing monitor, though it only tracks ten videos.\n\nThe soft spots are not minor. The arithmetic is internally inconsistent: the summary says 18% retrieved and 83% remained unavailable (which sums to 101%), and the category shares of 36% and 46% do not cleanly map onto the reported 18,961 IDs still not retrievable after re-querying. The 12.46% headline appears to be derived from the scraper-based estimate that 46% of the 70,239 initially missing videos were public but not available via the API, divided by 260,000. But the text explicitly says only 18,961 IDs were confirmed still not retrievable, which would give about 7.3% rather than 12.46%. If the 18,961 is only a subset because the individual re-query was interrupted, the paper never says so. That is load-bearing: the one-in-eight statistic is the main contribution, and it is not reproducible from the paper's own counts.\n\nAlso, the scraping ground truth is not documented—no classifier rules, no checks for false positives or negatives, no donor recruitment details. The paper does not release code or the full queried ID list, so independent verification is limited.\n\nThe conclusion that the API is incomplete is probably right, and the qualitative examples support it. But the current version does not support the specific quantitative claim. The paper deserves a serious referee because the topic matters for DSA compliance and platform accountability, but it needs heavy revision: fix the arithmetic, clearly separate observed re-query results from scraper-derived projections, document the scraping methodology, and release the data or code. I would send it to peer review with a request for major revision, not desk reject it. As it stands, I would not cite the 12.46% figure in my own work.","headline":"Real problems with TikTok's Research API, but the paper's own arithmetic makes its headline failure rate unreliable.","tokens_in":10473,"tokens_out":5630,"would_cite":false,"duration_ms":55201,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TikTok's Research API silently fails to return metadata for roughly one in eight donated public videos, with no error message explaining why.","keywords":["TikTok Research API","data donations","algorithm audits","Digital Services Act","data completeness","platform transparency","web scraping","advertising transparency"],"falsifier":"Run an independent crawler, from different IPs and without logging in, over the 70,239 initially missing IDs and compare its public/private classification with the paper's; if most of the videos the paper labels 'public but missing from the API' are actually inaccessible to the crawler, the 12.46% failure estimate collapses. Alternatively, if TikTok's Research API begins returning metadata for the dashboard's 10 persistently missing videos and the rate drops below one in eight on a fresh donated sample, the paper's conclusion would no longer hold.","tokens_in":9400,"feed_emoji":"📊","tokens_out":5492,"duration_ms":54740,"temperature":0.7,"pith_summary":"After Europe's Digital Services Act pushed TikTok to open its Research API to independent researchers, this audit tried to use it the way data-donation studies do: take a list of roughly 260,000 TikTok URLs donated by users and look up each video's metadata through the API. About one in eight videos could not be retrieved at all, and the API gave no error message saying whether the video was deleted, private, from a Canadian account, or something else. By scraping TikTok directly, the authors found that 62.7% of the unretrievable posts were actually still public on the platform, and they identified recurring categories of missing content: TikTok's own official videos, advertisements, and about 1% of creators whose videos never appear in the API. The paper's conclusion is that the Research API in its current state cannot serve research that needs complete, consistent, up-to-date data, and that researchers need scraping as an independent check.","feed_headline":"Research API omits 1 in 8 public TikTok videos","feed_subtitle":"Audit of 260,000 donated URLs finds public videos missing from official data with no error message.","key_machinery":"The load-bearing mechanism is the data-donation pipeline: a donated list of TikTok URLs is fed to the Research API's video endpoint in batches of 100, and any ID that returns no metadata is then checked by scraping the public TikTok site to decide whether it is deleted or private or public-but-missing. The key quantities are the initial failure set of 70,239 IDs, the eventual recovery of 12,373 IDs through individual re-querying, and the final classification of the remaining 18,961 IDs into categories (Canada, advertisements, TikTok-authorized videos, accounts excluded wholesale, and an unexplained remainder). The analysis also documents a 'corrupted batch' effect in which one unavailable ID in a batch makes other public IDs in the same batch appear unavailable, which is itself an API inconsistency.","core_discovery":"The central claim is that TikTok's Research API has a completeness and consistency problem, not merely a documentation gap. When 260,000 donated TikTok URLs were queried, 70,239 returned no metadata; after repeated individual queries and scraping, 18% of those were recovered, 36% had been deleted or made private, and 46% were confirmed public yet absent from the API. Within that public-but-missing share, Canadian creators (a documented restriction) and advertisements each accounted for a similar share, TikTok's own official videos were consistently absent, and 163 accounts had none of their videos accessible even though most were public. Because the API's error responses do not distinguish these causes, the authors conclude that a data-donation pipeline built on the API will silently drop roughly one in eight donated videos.","pith_inferences":["Because ads and TikTok's own company videos are overrepresented among the missing content, studies of political advertising and platform self-promotion that rely solely on the API will undercount exactly the content regulators care about under the DSA.","The 'corrupted batch' effect implies that non-retrieval is not always a property of the individual video; failure rates may depend on request composition, so even re-running the same query can change results.","The 12.46% figure is computed from one donated URL collection; extending the same scraping-plus-API check to random samples of For You Page content in multiple EU countries would show whether the failure rate is a stable property of the API or an artifact of that donor group."],"forward_implications":["Data-donation studies of TikTok will silently lose roughly 12.46% of donated videos, and because the loss is concentrated in ads, TikTok's own content, and certain creators, it is not a random sample of the missing content.","API-only research cannot distinguish 'the creator deleted this' from 'TikTok is withholding a public video', so moderation and removal analyses built on error codes will be unreliable.","Any finding that relies on the Research API's completeness, such as prevalence of ads or creator-level comparisons, needs an external scraping check before it can be interpreted.","The public dashboard makes the failure reproducible and observable over time, allowing researchers to check whether a problem they see is unique to their account or systemic."],"supporting_citations":[{"why":"Supplies the five data-quality dimensions (accuracy, completeness, timeliness, consistency, validity) that frame the audit.","marker":"Daikeler et al. (2024)"},{"why":"Documents TikTok's own statement that new videos take up to 48 hours to appear, the timeliness baseline the paper tests against.","marker":"TikTok Research API FAQ"},{"why":"Documents the API's exclusions of creators under 18, videos outside US/Europe/Rest of World, and Canadian videos, the known limitations the paper subtracts from its unexplained remainder.","marker":"TikTok Research API codebook"},{"why":"The page the paper says confusingly lists Canadian videos as an example, showing that even documentation of the known restrictions is inconsistent.","marker":"TikTok Research API Getting Started"}],"fun_headline_variants":["TikTok API silently drops 1 in 8 videos","Public videos vanish from TikTok research API","TikTok data gap: 1 in 8 videos unaccounted","Why TikTok's API hides public video metadata","TikTok research API omits 1 in 8 public videos"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The audit treats the authors' scraper's view of TikTok, what the public site shows from their vantage point, as the ground truth for whether a video is public, and treats the donated 260,000-URL list as representative of what data-donation research encounters.","fun_headline_variants_meta":{"raw":{"variants":["TikTok API silently drops 1 in 8 videos","Public videos vanish from TikTok research API","TikTok data gap: 1 in 8 videos unaccounted","Why TikTok's API hides public video metadata","TikTok research API omits 1 in 8 public videos"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000232,"raw_usage":{"total_tokens":1478,"prompt_tokens":921,"completion_tokens":557,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":475}},"tokens_in":537,"tokens_out":557,"duration_ms":6387,"temperature":1.0,"reasoning_tokens":475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:41:16.099133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an independent crawler, from different IPs and without logging in, over the 70,239 initially missing IDs and compare its public/private classification with the paper's; if most of the videos the paper labels 'public but missing from the API' are actually inaccessible to the crawler, the 12.46% failure estimate collapses. Alternatively, if TikTok's Research API begins returning metadata for the dashboard's 10 persistently missing videos and the rate drops below one in eight on a fresh donated sample, the paper's conclusion would no longer hold.","supporting_citations":[],"review_version":1}