Pith. sign in

REVIEW 36 references

Elephant in the Room: Dissecting and Reflecting on the Evolution of Online Social Network Research

T0 review · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that the only holistic dataset of OSN research reveals a field concentrated on Twitter and out of step with real-world platform use.

desk verdict First holistic OSN meta-analysis with a genuinely useful public dataset; sampling-frame bias is real but the qualitative findings are likely robust — send to review. read the letter →

arxiv 2411.13681 v2 pith:TCW5ZBQB submitted 2024-11-20 cs.SI

classification cs.SI
keywords OnlineSocialNetworksMinerva-OSNMeta-analysisTopicmodelingResearchtrendsDataaccesspoliciesTwitterprevalenceExpertsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish, for the first time, a holistic quantitative picture of all academic research on online social networks (OSNs) since 2006. It builds a public dataset, Minerva-OSN, of 13,842 peer-reviewed papers drawn from over a million candidates, and argues that the literature is lopsided: only 91 of 296 platforms have ever been studied, Twitter is the dominant subject since 2012 despite its declining popularity, and research output began dropping around 2018. The paper also argues that data-access policies of OSN owners are a root cause, and supports this with a manual review of eight platforms' APIs and a survey of 50 experienced researchers. If these claims hold, the community has the first evidence that OSN research is not representative of the platforms that people actually use, and a released dataset to test and extend that conclusion.

What carries the argument

Minerva-OSN, a curated dataset of 13,842 OSN papers, is the load-bearing object. Its construction has three stages: selecting 135 venues (top Google Scholar venues plus Lecture Notes in Computer Science and ACM proceedings), screening over one million abstracts with a heuristic that keeps a paper only if the abstract names at least one of 296 OSNs, and manual checks for homonyms and false positives. On top of this, a BERTopic pipeline, using sentence embeddings, UMAP, and HDBSCAN, assigns 17 topics, with OSN names replaced by a neutral tag to avoid platform bias; a tri-party human review of 170 abstracts validates the topic labels, achieving 72.5% first-choice agreement. The dataset and pipeline carry every subsequent claim about prevalence, topics, authors, and popularity mismatch.

What would settle it

Re-run the pipeline on a sample of Scopus-indexed health and psychology venues not in the 135-venue list, and on abstracts that mention 'social media' without naming a platform; if the Twitter share drops substantially or the number of studied OSN rises well above 91, the paper's descriptive claims are artifacts of the selection frame.

Watch

Extended reading notes

Core claim

The paper's central discovery is that the body of OSN research is concentrated on a few platforms, above all Twitter, and that this concentration has detached research from real-world platform usage. Among 296 OSNs considered, only 91 have been investigated; Twitter accounts for 5,248 of 13,842 papers, and it has been the most studied OSN since 2012 even though its website-traffic rank has fallen. Only 82 papers concern TikTok despite it having more than a billion users and a top-10 ranking. The paper connects this skew to data-access policies: platforms with free or cheap APIs, notably Twitter before April 2023, attracted research, while restrictive or expensive access deterred it. A survey of 50 researchers finds that 78% never collaborated with an OSN and that most call data access and reproducibility difficult. The conclusion is that OSN research carries nonrepresentative bias, with consequences for topics tied to younger demographics.

Load-bearing premise

The whole analysis depends on the selection rule that a paper is about an OSN only if its abstract names at least one OSN; papers that study online social networks without spelling out a platform, and venues outside the 135 chosen, are invisible to the dataset.

Editorial extensions

If this is right

  • Minerva-OSN gives researchers a quantitative baseline to measure whether future OSN research broadens beyond a handful of platforms.
  • The paper's evidence makes the case that restrictive data-access policies have a measurable effect on what gets studied.
  • Topical gaps follow from the skew: issues affecting younger demographics are concentrated on platforms like TikTok and Instagram that the literature has yet to cover adequately.
  • The 17-topic model and its tri-party validation give future literature analyses a reusable procedure instead of an ad hoc manual review.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the selection heuristic were widened to abstracts using generic phrases like 'social media' without naming a platform, method-only and cross-platform papers would likely enter the dataset; the reported Twitter share could then shrink, so the dominance number is partly a property of the search rule.
  • The 2018 decline in US and EEA output could be checked against a matched corpus of non-OSN papers from the same venues; if those also decline, the GDPR explanation is weaker than a general publication trend.
  • A testable projection arises: the 2023 API paywalls at Twitter and Reddit should push future OSN research toward platforms with open access, such as Mastodon or Wikipedia, and reduce Twitter's share in follow-up versions of Minerva-OSN.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the meta-analysis tabulates an explicitly scoped dataset, and its coverage limitations are acknowledged rather than concealed.

full rationale

The paper's derivation chain is a systematic literature meta-analysis: it constructs Minerva-OSN from a stated venue list and an abstract-mention heuristic, then reports descriptive statistics over that dataset. The counts that feed RQ1-RQ5 (e.g., 91 of 296 OSN investigated, Twitter predominance) are tabulations of the inclusion criteria, not outputs of a fitted model or of a self-citation chain. The heuristic's blind spot for abstracts that do not name a platform is explicitly disclosed in Limitations ('our heuristic may have not captured papers that did not mention any specific OSN in the abstract'), which is a coverage limitation rather than a circular reduction. The topic-model pipeline is validated by a blinded tri-party human review, and the expert survey collects independent opinions about data-access difficulty; neither is used to define the quantities it is said to support. The only self-citations (e.g., Sartori et al. 2023) appear in related-work context and are not load-bearing for any central claim. No equation, fitted parameter, or imported uniqueness theorem is used to force a result, so no circular step is exhibited.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new physical or theoretical entities. Its free parameters are methodological choices: the threshold of OSN mention, the venue selection, the topic model configuration, and the survey design. These choices shape all reported statistics, and several are acknowledged limitations in the paper itself. No entity with a falsifiable outside handle is introduced.

free parameters (3)
  • OSN mention threshold for paper inclusion = at least one OSN name in the abstract
    The heuristic defining Minerva-OSN is a choice that determines all downstream statistics; it is not derived from any external criterion.
  • Topic model hyperparameters = not fully specified in the paper
    BERTopic, UMAP, HDBSCAN, and bge-small-en-v1.5 settings are described at a high level, but the exact parameters are not reported in the manuscript and would affect topic assignments.
  • Venue selection criteria = 135 venues from Google Scholar top venues, LNCS, and ACM ICPS
    The choice of venues is a modeling decision that defines the population of papers; it is a hand-chosen inclusion criterion, not a derived quantity.
assumptions (3)
  • domain assumption Scopus metadata is complete and accurate for the selected venues
    The entire dataset depends on the completeness of Scopus records, which the authors acknowledge may be incomplete for 2023.
  • domain assumption Google Scholar's top 20 venues per subcategory, plus LNCS and ACM ICPS, constitute a representative sample of OSN research venues
    This assumption defines the scope of Minerva-OSN and is not independently verified; it could exclude important venues outside computer science.
  • domain assumption Mentioning an OSN name in the abstract is a valid proxy for being an OSN research paper
    The authors state this as the core heuristic and acknowledge it may miss relevant papers; it is a domain modeling assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Elephant in the Room: Dissecting and Reflecting on the Evolution of Online Social Network Research." pith.science (2026). https://pith.science/paper/TCW5ZBQB

@misc{pith2026241113681,
  author       = {Pith},
  title        = {Pith review of: Elephant in the Room: Dissecting and Reflecting on the Evolution of Online Social Network Research},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TCW5ZBQB}},
  note         = {Machine review of arXiv:2411.13681}
}
read the original abstract

Billions of individuals engage with Online Social Networks (OSN) daily. The owners of OSN try to meet the demands of their end-users while complying with business necessities. Such necessities may, however, lead to the adoption of restrictive data access policies that hinder research activities from "external" scientists -- who may, in turn, resort to other means (e.g., rely on static datasets) for their studies. Given the abundance of literature on OSN, we -- as academics -- should take a step back and reflect on what we have done so far, after having written thousands of papers on OSN. This is the first paper that provides a holistic outlook to the entire body of research that focused on OSN -- since the seminal work by Acquisti and Gross (2006). First, we search through over 1 million peer-reviewed publications, and derive 13,842 papers that focus on OSN: we organize the metadata of these works in the Minerva-OSN dataset, the first of its kind -- which we publicly release. Next, by analyzing Minerva-OSN, we provide factual evidence elucidating trends and aspects that deserve to be brought to light, such as the predominant focus on Twitter or the difficulty in obtaining OSN data. Finally, as a constructive step to guide future research, we carry out an expert survey (n=50) with established scientists in this field, and coalesce suggestions to improve the status quo such as an increased involvement of OSN owners. Our findings should inspire a reflection to "rescue" research on OSN. Doing so would improve the overall OSN ecosystem, benefiting both their owners and end-users and, hence, our society.

Figures

Figures reproduced from arXiv: 2411.13681 by the authors.

Figure 1
Figure 1. OSN prevalence over time. We highlight the dis￾crepancy between OSN investigated in research and their real-world popularity (website ranking based on Alexa [25]). release. Next, through a meta-analysis, we investigate our Minerva-OSN dataset across four dimensions: OSN preva￾lence, wherein we show that some OSN are significantly more studied than others (see [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distribution of topics in Minerva-OSN. After processing the 13,842 papers in dataset, our topic model identified 17 topics which described the abstracts of 83.9% of the papers in Minerva-OSN (topic description in [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Yearly distribution of the most prevalent OSN. [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Topics over Time (for 8 topics). The y-axis reports the number of papers on each topic. Thick lines show the trend of that specific topic, thin lines show all the others. RQ2: Analysis of the Authors of OSN Papers What are the predominant patterns among authors researc…
Figure 5
Figure 5. Figure 5: Authors’ affiliation country (yearly distribution). [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 7
Figure 7. Figure 7: Left: papers from EEA/non-EEA researchers. [PITH_FULL_IMAGE:figures/full_fig_p007_7.png]
Figure 8
Figure 8. Figure 8: Papers (related to Twitter, Facebook, Reddit, TikTok) authored by EEA and non-EEA researchers over time. The y-axis of each subfigure has a different scale because we compare trends between EEA and other countries for each OSN (we do not aim at carrying out cross-platf…
Figure 9
Figure 9. Figure 9: Researchers preference in data gathering: ideal vs. real world. We indicate preference (1 is low, 5 is high) to rely on open source datasets (“Open”), creating a new dataset through official APIs (“API”) or by scraping (“scrape”), or via a direct channel with the OSN (…
Figure 10
Figure 10. Figure 10: BERTopic visualization. Topic Name Description Size Multimedia Retrieval and Tagging It focuses on developing methods to efficiently search and categorize visual content, such as images and videos, often leveraging techniques like face recognition and object detection…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

36 extracted references · 27 canonical work pages

  1. [1]

    For most authors... (a) Would answering this research question advance sci- ence without violating social contracts, such as violat- ing privacy norms, perpetuating unfair profiling, exac- erbating the socio-economic divide, or implying disre- spect to societies or cultures? Yes (b) Do your main claims in the abstract and introduction accurately reflect t...

  2. [2]

    Additionally, if your study involves hypotheses testing... (a) Did you clearly state the assumptions underlying all theoretical results? NA (b) Have you provided justifications for all theoretical re- sults? NA (c) Did you discuss competing hypotheses or theories that might challenge or complement your theoretical re- sults? NA (d) Have you considered alt...

  3. [3]

    (a) Did you state the full set of assumptions of all theoret- ical results? NA (b) Did you include complete proofs of all theoretical re- sults? NA

    Additionally, if you are including theoretical proofs... (a) Did you state the full set of assumptions of all theoret- ical results? NA (b) Did you include complete proofs of all theoretical re- sults? NA

  4. [4]

    Additionally, if you ran machine learning experiments... (a) Did you include the code, data, and instructions needed to reproduce the main experimental results (ei- ther in the supplemental material or as a URL)? Yes, we provided a URL with source code and data. (b) Did you specify all the training details (e.g., data splits, hyperparameters, how they wer...

  5. [5]

    (a) If your work uses existing assets, did you cite the cre- ators? NA, as we only utilized the dataset we collected

    Additionally, if you are using existing assets (e.g., code, data, models) or curating/releasing new assets,without compromising anonymity... (a) If your work uses existing assets, did you cite the cre- ators? NA, as we only utilized the dataset we collected. (b) Did you mention the license of the assets? NA (c) Did you include any new assets in the supple...

  6. [6]

    (a) Did you include the full text of instructions given to participants and screenshots? Yes, in the online repos- itory we share a copy of the survey

    Additionally, if you used crowdsourcing or conducted research with human subjects,without compromising anonymity... (a) Did you include the full text of instructions given to participants and screenshots? Yes, in the online repos- itory we share a copy of the survey. (b) Did you describe any potential participant risks, with mentions of Institutional Revi...

  7. [8]

    https://anonymous.4open.science/r/Minerva-OSN-3841

  8. [9]

    https://huggingface.co/spaces/mteb/leaderboard

Show all 36 references
  1. [10]

    https://blog.flickr.net/en/2018/11/01/changing-flickr- free-accounts-1000-photos/

  2. [11]

    https://www.linkedin.com/pulse/important-linkedin- statistics-data-trends-oleksii-bondar-pqlie/

  3. [14]

    https://web.archive.org/web/20240303190109/https: //developers.google.com/youtube/v3/docs/search/list

  4. [18]

    https://web.archive.org/web/20231204034413/https: //steamcommunity.com/dev

  5. [20]

    https://web.archive.org/web/20231130070701/https: //developer.twitter.com/en/docs/twitter-api

  6. [21]

    https://thesmallbusinessblog.net/how-many-people-use- yelp/

  7. [22]

    https://web.archive.org/web/20231128205404/https: //developers.tiktok.com/products/research-api/

  8. [23]

    https://web.archive.org/web/20231210080026/https: //fort.fb.com/researcher-apis

  9. [24]

    https://web.archive.org/web/20240512003912/https: //www.reddit.com/r/AskAcademia/comments/1b32i9q/ accessing reddit data for academic purposes/?rdt= 37952

  10. [25]

    https://www.statista.com/statistics/272014/global- social-networks-ranked-by-number-of-users/

  11. [26]

    https://en.wikipedia.org/wiki/Wikipedia:Wikipedians

  12. [27]

    https://photutorial.com/flickr-statistics/

  13. [28]

    https://www.statista.com/statistics/1337525/us- distribution-leading-social-media-platforms-by-age- group/

  14. [29]

    https://business.quora.com/resources/reach-over-400- million-monthly-unique-visitors-on-quora/

  15. [30]

    https://www.statista.com/forecasts/1309791/reddit-mau- worldwide

  16. [31]

    https://perspectiveapi.com/research/

  17. [32]

    https://web.archive.org/web/*/https://www.alexa.com/ topsites

  18. [33]

    https://en.wikipedia.org/wiki/List of social networking services

  19. [34]

    https://en.wikipedia.org/wiki/Alt-tech

  20. [35]

    https://usesignhouse.com/blog/stack-overflow-stats/#: ∼:text=How%20many%20registered%20users%20of% 20Stack%20Overflow%20are%20there%3F,-Want% 20a%20link&text=As%20of%20November%202022% 2C%20there%20are%20more%20than%20100% 20million,and%2023%20million%20registered% 20users

  21. [36]

    https://radar.cloudflare.com/domains

  22. [37]

    https://www.go-fair.org/fair-principles/

  23. [38]

    https://scholar.google.com/citations?view op= top venues&hl=en&vq=eng

  24. [39]

    https://www.statista.com/statistics/1314183/youtube- shorts-performance-worldwide/

  25. [40]

    https://www.statista.com/statistics/795303/china-mau- of-sina-weibo/

  26. [41]

    https://www.statista.com/statistics/234038/telegram- messenger-mau-users/

  27. [43]

    com/2023/04/18/technology/reddit-ai-openai- google.html

    https://web.archive.org/web/2/https://www.nytimes. com/2023/04/18/technology/reddit-ai-openai- google.html

  28. [2019]

    McInnes, L.; Healy, J.; and Melville, J

    An analysis of the consequences of the general data protec- tion regulation on social network research.TSC. McInnes, L.; Healy, J.; and Melville, J. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv:1802.03426. O’Connor, C.; and Joffe, H....

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.