{"id":"f7c2ded6-a7a2-4557-a294-86b058f7c908","arxiv_id":"2505.07606","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 29 Mastodon studies finds that researchers rarely engage with instance-level data policies, prompting calls for structural fixes to research ethics on the Fediverse.","lead":"This paper reviews 29 academic studies that used Mastodon data and finds that most researchers paid little attention to the rules of individual Mastodon servers, even when those rules explicitly ban data collection. It argues for new ethical guidelines and technical tools so that research on decentralized social networks respects community policies.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Temporal mismatch makes the non-adherence claim retrospective: policies checked on March 27, 2025 may postdate the studies' data collection windows.","rationale":"The reader's verdict of CONDITIONAL with moderate confidence is well calibrated, and the reader's weakest_assumption correctly identifies the most load-bearing issue: the March 27, 2025 policy snapshot is applied to studies whose data collection occurred earlier. I agree with this concern and find no other more serious flaw. The paper is transparent about many limitations—keyword-search proxies, OpenAlex indexing gaps, and the non-binding nature of robots.txt—but it does not address the temporal anchoring problem in its Limitations section. The central claim would hold if instance policies relevant to the surveyed studies were stable over time, but the paper provides no evidence of stability and even documents one domain reassignment. A concrete Wayback Machine check can settle whether the specific prohibitions found in 2025 existed during each study's collection window. If the check fails, the abstract's 'revealing limited adherence' is too strong and should be softened to something like 'revealing limited current alignment with instance policies.' If it succeeds, the finding is materially strengthened. In either case the conditional verdict remains appropriate: the paper's recommendations do not depend on the historical finding, and the overall systematic-review contribution is sound. No formal verification or shared search artifacts exist, but the full list of surveyed works is provided, which aids reproducibility. I do not see a basis for rejection or for unconditional acceptance; the verdict should stay CONDITIONAL pending the temporal check.","tokens_in":11709,"tokens_out":1647,"duration_ms":20556,"concrete_test":"For each of the 29 surveyed works, extract the Mastodon instances used and the stated data collection dates from the paper or its appendices; then retrieve archived versions of each instance's rules and privacy policies from the Wayback Machine (web.archive.org) at timestamps covering the start and end of that study's collection window. Re-run the adherence classification for the 'at least two instances' prohibition claim and for the four datasets with allegedly incompatible licenses using only policy text archived before or during collection. If the historical policy text contains the same prohibitions, the temporal concern is resolved and the central claim is supported; if not, the abstract and Results should be revised to report current-policy mismatches rather than historical non-adherence.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that published Mastodon research shows 'limited adherence to instance-level policies' depends on comparing each study's practices against instance rules, but the paper checks policies at a single date. The Results section states that 'at least two of the referenced instances prohibit data collection without user consent (as of March 27, 2025)' and evaluates four published datasets' licenses against instance rules 'as of' that same date. Most surveyed studies collected data years earlier, and Mastodon instances can change, delete, or add policies without notice; the paper itself notes that mnm.social's domain was reassigned by March 27, 2025, demonstrating policy instability. Without evidence that each relevant rule existed during the corresponding study's data collection period, findings of 'non-adherence' or 'inappropriate licenses' describe a current mismatch rather than a failure of researchers' choices at the time. The abstract's stronger phrasing—'revealing limited adherence'—thus overstates what the single-date comparison can establish. The paper's own Limitations section acknowledges keyword-based screening and OpenAlex indexing issues, but not this temporal dependency, which is the most load-bearing unaddressed assumption. If historical policies differed, the headline result could reverse for the specific 'at least two instances' and license incompatibilities that drive the conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a systematic literature review of 29 academic works that collected user-generated data from Mastodon. It examines how researchers handled instance-specific policies, privacy protections, and data publication, concluding that most works show limited engagement with instance-level policies despite mentioning them, and that some published datasets use licenses incompatible with instance rules. The paper also proposes recommendations for researchers, ethics committees, software developers, and instance administrators.","tokens_in":11894,"tokens_out":4088,"duration_ms":39343,"significance":"If the findings hold, this is a valuable contribution to social media research ethics, extending earlier work on Reddit to decentralized platforms. The paper's strengths include a transparent search protocol, explicit inclusion/exclusion criteria, a full list of surveyed works, and attention to the 2019 Zignani retraction case. The main limitations concern the temporal validity of policy checks and the inferential gap from keyword mentions to 'awareness.'","major_comments":[{"comment":"The comparison of research practices against instance policies relies on policies as they stood on March 27, 2025, while most of the 29 works collected data between 2018 and 2025. The paper states that 'at least two of the referenced instances prohibit data collection without user consent (as of March 27, 2025)' and judges dataset licenses 'as of' the same date. Because instance policies can change without notice—the paper itself notes that mnm.social's domain was reassigned by that date—these checks establish a current mismatch, not necessarily a failure by researchers at the time of data collection. The Limitations section acknowledges keyword-based screening and OpenAlex indexing issues but not this temporal dependency. To support the headline claim of 'limited adherence,' the authors should either verify policies via archived versions (e.g., Wayback Machine) for the relevant data-collection periods or explicitly re-frame the findings as an assessment of current policy compatibility.","section":"Results (policies and dataset licenses); Limitations"},{"comment":"The conclusion states that 'seven works have published their data, with four applying inappropriate licenses.' In the Results, however, only two of those datasets are directly shown to be incompatible with the rules of the single identified instance; for the other two, the text says it is 'likely' they include content from noncompliant instances. The shift from 'likely' to a definite count of four overstates the evidence and should be corrected in the abstract and conclusion, or the analysis should be extended to verify those two datasets.","section":"Results (published data licenses); Conclusion and Discussion"},{"comment":"The claim that researchers have 'general awareness' of instance policies is inferred from a keyword search for 'polic-', 'rule', and 'terms' in the 29 works. Mere occurrence of these terms does not demonstrate awareness of the specific policies of the instances from which data were collected; the paper itself notes that 11 works acknowledged governance documents 'without clear implications for their methods.' The Limitations section concedes the keyword-based screening is coarse. The operationalization of 'awareness' should be described as such, and the claim in the abstract should be tempered to indicate that papers mention policies, not that researchers are aware of their content.","section":"Results (policy keyword search); Limitations"}],"minor_comments":[{"comment":"The word 'ressources' should be 'resources', and the reference to 'OpenAlex.org 2025' appears with an inconsistent period in the reference list.","section":"Data and Methods"},{"comment":"The appendix lists the 29 works but does not provide a per-paper coding table; including a table with each work's data collection method, instance selection, policy mention, and data publication status would improve transparency and reproducibility.","section":"Appendix"},{"comment":"Table 1 contains a footnote marker ('3') that is not explained in the table or the surrounding text; please clarify what this footnote refers to.","section":"Table 1"},{"comment":"The phrase 'a dummy instance' in the instance selection paragraph is unclear; please define what constitutes a dummy instance and why a study might use one.","section":"Results"},{"comment":"Some references are incomplete or informally cited, such as 'Cathleen O’Grady 2025' appearing in the text without a full reference entry, and several arXiv papers lacking version or DOI information.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope well, but the authors should be asked to provide their coding instrument and per-paper evidence for the policy checks, especially the archived policy data. The temporal mismatch is the most serious issue; if archived policies can be obtained, the paper could become a strong contribution. I would also encourage the authors to report the inter-coder reliability of their qualitative judgments, as the current text presents them as straightforward observations."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague – quick take on arXiv:2505.07606. This is the first systematic review I've seen of how Mastodon studies actually handle instance policies, data sharing, anonymization, and third-party transfers. The corpus of 29 works is small but carefully assembled, and the coding scheme is transparent enough to follow. The findings – most papers mention policies only vaguely, seven share data, four use licenses that clash with instance rules, and several transfer data to external APIs – line up with the broader Reddit ethics literature and give the field something concrete to argue with. The recommendations (machine-readable versioned instance rules, ethics-board guidance, API documentation changes) follow reasonably from the evidence.\n\nThe soft spots are real but not fatal. The main one is the temporal mismatch. Instance policies are checked as of March 27, 2025, while most studies collected data years earlier. Mastodon policies change, and the paper itself notes mnm.social was reassigned by that date. So the claim that 'at least two of the referenced instances prohibit data collection without user consent (as of March 27, 2025)' – and the license incompatibility judgments – describe a current mismatch, not necessarily what researchers faced during collection. The abstract's 'limited adherence' overstates this. The Limitations section does not mention this dependency, and it should.\n\nTwo smaller issues: 'awareness' is inferred from keyword mentions (polic-, rule, terms), which is a proxy, and the seven papers added from the author's private library are a bit of a black box – fine for transparency, but it means the corpus isn't fully reproducible from OpenAlex alone. Inter-rater reliability is not reported, though with two authors coding, it would have been cheap to add.\n\nThat said, the core pattern – shallow engagement with instance-level governance in Mastodon research – is likely robust. Most of the evidence doesn't depend on the exact policy date; it depends on what researchers wrote in their own papers. The citation pattern looks fine, and the discussion of the Zignani retraction is appropriately contextualized.\n\nWho's this for: people doing social media research ethics, especially on decentralized platforms, and anyone reviewing such work. It merits a serious referee – send it out. I'd want the temporal issue addressed before publication, but the paper is a useful contribution to the niche.\n\nRecommendation: accept with revisions, as long as the abstract and claims are aligned with the evidence and the policy-check date limitation is acknowledged.","headline":"Useful first systematic review of Mastodon research ethics, but the temporal mismatch between policy checks and data collection weakens the strongest claim in the abstract.","tokens_in":12422,"tokens_out":1776,"would_cite":true,"duration_ms":16836,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Mastodon researchers rarely follow the rules of the instances they collect data from, even when they know the rules exist.","keywords":["Mastodon","Fediverse","instance policies","data ethics","systematic literature review","terms of service","social media research","research ethics"],"falsifier":"For each of the 29 studies, pull archived copies of the referenced instances' policies dated to that study's data-collection window and compare them with the March 27, 2025 versions. If none of the instances prohibited data collection or redistribution at the relevant time, the central finding of non-adherence collapses; the current-policy comparison would then measure policy change rather than researcher neglect.","tokens_in":11483,"feed_emoji":"🐘","tokens_out":7016,"duration_ms":62207,"temperature":0.7,"pith_summary":"This paper tries to establish that the ethical gap long documented in social media research carries over to decentralized platforms: studies that collect Mastodon data rarely engage with the policies of the individual instances they draw from, even though those policies often say such collection is not allowed. The claim rests on a systematic review of 29 works that used Mastodon as a data source, checking how they selected instances, whether they mentioned instance rules, how they handled privacy, and whether they published data. The stakes are concrete: if the claim is right, then a systematic body of published research and public datasets rests on practices that violate the communities' own rules, at a time when researchers are turning to Mastodon precisely because centralized platforms closed their APIs. The paper argues for treating each instance as a distinct community with enforceable ethical norms, not as an undifferentiated data source.","feed_headline":"Most Mastodon research ignores server rules","feed_subtitle":"Review of 29 works finds few check instance policies; some release data under incompatible licenses.","key_machinery":"The policy-adherence audit: a coding scheme that, for each of the 29 works, records the data-collection method, the number and selection of instances, whether instance policies are mentioned, whether data were published and under which license, and whether anonymization is claimed or effective. The audit is what converts individual cases into a pattern, and its comparison of published licenses against current instance rules is the mechanism that exposes the incompatibilities. The review also uses keyword-based content screening and snowball sampling to build the corpus, but the load-bearing step is this systematic comparison between what the studies say they did and what the instances' policies allow.","core_discovery":"The paper's central claim is that in Mastodon research, awareness of instance-level governance does not translate into adherence. Of 29 surveyed works, 17 touched on policies, but most acknowledged instance-specific governance documents without drawing clear implications for their own methods; only one conducted a manual review of the terms of service for the instances it used. At least two of the instances referenced in the surveyed works currently forbid data collection without user consent, and one explicitly bans use for AI training. Seven works published Mastodon data, and four of those datasets were released under licenses that conflict with the source instances' rules. The paper also reports that data sharing is often done with anonymization that does not actually prevent re-identification, since toot content can be cross-referenced to identify users even when user IDs are obfuscated.","pith_inferences":["If the non-adherence is as systematic as this review suggests, downstream re-users of existing Mastodon datasets inherit the original violation; a practical extension would be a registry of instance-policy versions that researchers must consult before releasing data.","A natural replication is to run the same audit on other fediverse software such as Pleroma, Misskey, or Lemmy, to test whether the gap is specific to Mastodon or general to decentralized social media.","Because the review judges policies as of March 27, 2025, a longitudinal replication using archived policies from each study's collection window would show whether the neglect is stable or a recent phenomenon; the paper's own call for machine-readable rules points in this direction."],"forward_implications":["Researchers collecting Mastodon data should treat each instance's rules as a distinct ethical constraint, and ethics boards should ask which instances were used and what their policies permit.","Published Mastodon datasets may carry licenses their source instances do not allow, so dataset publishers may need to re-check, re-license, or withdraw existing data.","Tool and API documentation share responsibility: data-collection libraries that warn about instance policies or implement checks would address a structural cause of the gap.","The pattern matches earlier findings from Reddit research, suggesting this is a systemic issue for community-based platforms rather than a handful of negligent studies."],"supporting_citations":[{"why":"Large-scale overview of Reddit research showing only 14% mentioned ethics approval; provides the comparative baseline the paper's findings mirror.","marker":"Proferes et al. 2021"},{"why":"Qualitative follow-up on Reddit that frames ethics considerations as often procedural; source for the recommendation that communities be treated as participants.","marker":"Fiesler et al. 2024"},{"why":"Guide arguing the Mastodon API flattens complexity and that API access does not imply user consent; supplies the paper's core normative frame.","marker":"Roscam Abbing and Gehl 2024"},{"why":"Survey of instance rules finding 31 of 4,371 instances explicitly resist research or indexing; establishes that policy prohibitions are common.","marker":"Wähner et al. 2024"},{"why":"The retracted study that motivates the review; its justification and removal frame the research questions.","marker":"Zignani et al. 2019c"},{"why":"Open letter critiquing the retracted study over terms-of-service and privacy violations; documents community expectations.","marker":"Administrators, Scholars and Users 2020"},{"why":"Official documentation with a 'Playing with public data' section; cited as a structural factor that can mislead researchers into thinking public posts are freely usable.","marker":"Mastodon API 2024"},{"why":"Evidence that users do not necessarily expect their public posts to become research data; supports the ethical weight of instance rules.","marker":"Fiesler and Proferes 2018"}],"fun_headline_variants":["Only 1 in 29 Mastodon studies checks server rules","Mastodon research data often violates instance policies","Mastodon research licenses clash with server rules","Anonymization fails in Mastodon research data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review judges adherence using instance policies as they stood on March 27, 2025, while most surveyed studies collected data earlier, so if servers changed their rules after collection, the non-adherence findings would describe policy drift, not the choices researchers actually faced.","fun_headline_variants_meta":{"raw":{"variants":["Only 1 in 29 Mastodon studies checks server rules","Mastodon research data often violates instance policies","Mastodon research licenses clash with server rules","Anonymization fails in Mastodon research data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000588,"raw_usage":{"total_tokens":2671,"prompt_tokens":768,"completion_tokens":1903,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":384,"completion_tokens_details":{"reasoning_tokens":1841}},"tokens_in":384,"tokens_out":1903,"duration_ms":11802,"temperature":1.0,"reasoning_tokens":1841,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:11:45.740558+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"For each of the 29 studies, pull archived copies of the referenced instances' policies dated to that study's data-collection window and compare them with the March 27, 2025 versions. If none of the instances prohibited data collection or redistribution at the relevant time, the central finding of non-adherence collapses; the current-policy comparison would then measure policy change rather than researcher neglect.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large-scale overview of Reddit research showing only 14% mentioned ethics approval; provides the comparative baseline the paper's findings mirror."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Qualitative follow-up on Reddit that frames ethics considerations as often procedural; source for the recommendation that communities be treated as participants."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Guide arguing the Mastodon API flattens complexity and that API access does not imply user consent; supplies the paper's core normative frame."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Open letter critiquing the retracted study over terms-of-service and privacy violations; documents community expectations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Official documentation with a 'Playing with public data' section; cited as a structural factor that can mislead researchers into thinking public posts are freely usable."},{"cited_title":"Participant","cited_arxiv_id":null,"evidence_quote":"Evidence that users do not necessarily expect their public posts to become research data; supports the ethical weight of instance rules."}],"review_version":1}