{"id":"d9c3649e-facf-48ef-ab5d-6cc217212678","arxiv_id":"2412.16383","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Mastodon users who joined after the Twitter acquisition form ego networks compatible with Dunbar's model, with 4-5 layers and a scaling ratio near 3, though the networks are still young.","lead":"This paper analyzes the social circles of nearly 2,000 Mastodon users who joined after Elon Musk acquired Twitter, using replies, mentions, and boosts to map each user's ego network. It finds these networks fit Dunbar's layered model with 4 to 5 circles and a scaling ratio near 3, but with smaller outer layers, a sign of a young social network.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline scaling-ratio claim is not supported by the paper's own Table VII: most consecutive-layer ratios are 1.8-2.9 or 4.4-5.9, not ~3, so 'compatible with the canonical Dunbar model' rests on a weaker pillar than stated.","rationale":"I read the paper as attempting to establish that Mastodon ego networks show the canonical Dunbar structure, specifically 4-5 circles and roughly 3x scaling between layers. The preponderance of 4 and 5 circles is plausible from Figure 7, and the activity analysis is a useful contribution. However, the scaling-ratio evidence is the weaker pillar of the compatibility claim. The mismatch is not simply a matter of sample representativeness; it is visible in the authors' own summary tables. If per-layer ratios are highly variable and the outer ratio is systematically 4-6, the phrase 'scaling ratio (~3) across circles' overstates the fit. The paper's caveat about a young network explains the outer layer, but it does not establish that the remaining inner ratios follow a single ~3 geometric progression. I would not reject the manuscript: the qualitative layered-structure finding and the useful activity analysis remain, and a re-analysis or re-framing could resolve the issue. The single-seed snowball concern raised by the reader is real, but it is secondary here because it affects generalizability, not the internal mismatch between the stated claim and the reported numbers. If the ratio evidence collapses, the compatibility claim fails regardless of seed choice. Hence the conditional status with a request for per-ego ratio distributions and confidence intervals is appropriate; my read does not move the verdict.","tokens_in":12920,"tokens_out":7585,"duration_ms":60681,"concrete_test":"Recompute from the raw timeline data the per-ego vector of circle sizes and per-ego consecutive ratios, then fit log(size_k) against layer index k for each ego and for all egos pooled. Report 95% confidence intervals for the per-transition slope (log-ratio) and use a bootstrap to test the null hypothesis that the geometric mean ratio equals 3. If, as Table VII suggests, most confidence intervals exclude 3 and the dispersion is large, the central claim should be revised to describe the observed young-network deviations rather than claiming compatibility with canonical ~3 scaling. This requires releasing or sharing the per-ego circle-size data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section I asserts that Mastodon ego networks are compatible with the canonical Dunbar model 'in terms of the number of circles ... and scaling ratio (~3) across circles.' The paper's own aggregate tables undercut this. From Table VI, rows for 4-circle and 5-circle egos give sizes roughly 1.4/3.7/10.6/45.8 and 1.2/2.8/5.8/15.4/77.0; the corresponding ratios in Table VII are 2.85, 2.95, 4.41 and 2.48, 2.13, 2.64, 4.75. For 6-circle egos the ratios are 2.27, 1.81, 1.85, 2.78, 5.91. (Table VII's headers 3/4, 4/5, and 6/5 appear to be typographical inversions of 4/3, 5/4, and 6/5.) Only a minority of the sixteen ratios is close to 3; the outermost-layer ratio is 4.4-5.9, and several inner ratios are near 2 rather than 3. The paper acknowledges the outer ratio is higher, but 'frequently close to 3' is an average-ratio gloss, not the per-layer pattern. Because the '~3 scaling between circles' is half of the compatibility claim, this is a load-bearing mismatch that would remain even with a perfectly representative sample.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes user-to-user directed interactions on Mastodon after the Twitter/X acquisition, using a snowball sample of roughly 2,000 egos from a single seed on mastodon.social. For the post-acquisition active user group, it constructs ego networks by clustering alter contact frequencies with the Meanshift algorithm and reports the number of layers, layer sizes, scaling ratios, and contact frequencies. The central claim is that Mastodon ego networks are compatible with the canonical Dunbar model, with a preponderance of 4 and 5 circles and a scaling ratio of about 3 across layers. The paper also interprets smaller outer layers and higher outer-layer ratios as signs of a 'young' social network. It concludes that Mastodon, with its open API, is a viable replacement for studying human social behavior.","tokens_in":13201,"tokens_out":3473,"duration_ms":31701,"significance":"If the central claim were fully supported, the paper would establish Mastodon as a convenient laboratory for validating the Dunbar ego-network model in a decentralized, open platform. The work is among the first to apply the Dunbar ego-network pipeline to Mastodon and contributes a large, publicly obtainable dataset of directed interactions. The methodology follows a well-established pipeline from the authors' prior work, which is a strength for comparability. However, the headline compatibility claim is currently overstated relative to the paper's own Table VII, and the single-seed snowball sample limits how far the results can be generalized to Mastodon as a whole. These issues are load-bearing for the paper's main conclusion, though they are addressable with a more measured claim and additional analysis.","major_comments":[{"comment":"The statement that 'the scaling ratio is frequently close to 3' is not supported by the data in Table VII. For 4-circle ego networks the ratios are 2.85, 2.95, and 4.41; for 5-circle networks they are 2.48, 2.13, 2.64, and 4.75; for 6-circle networks they are 2.27, 1.81, 1.85, 2.78, and 5.91. Only a minority of the entries are near 3, and the outermost-layer ratio is consistently 4.4-5.9. Because the '~3 scaling ratio' is half of the compatibility claim made in Section I, this mismatch is load-bearing. The authors should either revise the claim to state that inner-layer ratios are typically 1.8-3.0 while the outermost ratio is systematically larger in this young population, or provide a statistical test (e.g., confidence intervals for the mean ratios) showing that the ratios are consistent with 3 after accounting for variability.","section":"Section VII, Table VII"},{"comment":"The sample is a single snowball component starting from one 'random user highly active in the first years of Mastodon' on mastodon.social, with active alters prioritized and collection stopping at 2,000 users. The paper acknowledges that 'all collected users belong to the same connected component.' This design assumes that one component is representative of Mastodon as a whole, which is a strong assumption given prior evidence of instance-level heterogeneity (e.g., Zignani et al., ref [25], and La Cava et al.). Because the central claim is about Mastodon generally, the manuscript should either add a sensitivity analysis with multiple seeds from different instances and communities, or explicitly restrict the conclusions to the studied component and explain why the single-component generalization is still justified.","section":"Section IV"},{"comment":"The interpretation of 'young ego networks' rests on the premise that external layers take longer to stabilize, but the analysis does not test this explanation against alternatives. In particular, the preprocessing filters (alters contacted at least twice with annual frequency >1, and relationships lasting at least six months) will disproportionately censor the outer layers because those layers have low contact frequency and may include relationships shorter than six months within the ~1.5-year observation window. The observed smaller outer layers and larger outermost scaling ratio could be artifacts of these filters rather than a universal property of young networks. The authors should quantify how the filter thresholds affect the layer sizes and scaling ratios, for example by varying the six-month threshold and the annual frequency cutoff in a robustness check.","section":"Section VI-B and Section V-B"},{"comment":"All reported layer sizes, scaling ratios, and contact frequencies are simple averages over egos within each circle-count group. The paper does not provide standard deviations, confidence intervals, or a test for whether the observed values differ significantly from the canonical Dunbar values (1.5, 5, 15, 50, 150). Without such statistics, it is difficult to assess whether the observed differences (e.g., inner layers around 1.2-1.4 instead of 1.5, or outer layers around 77 instead of 150) are meaningful. Adding per-group distributions or error bars would materially strengthen the compatibility claim.","section":"Tables V-VII"}],"minor_comments":[{"comment":"The column headers '3/4', '4/5', and '6/5' appear to be typographical inversions of '4/3', '5/4', and '6/5'; if so, they should be corrected because the values in the table correspond to a ratio larger than 1 for the outermost layers.","section":"Table VII"},{"comment":"The condition 'Cij >= 2 and Fij > 1' is redundant in some timing configurations but not in others (e.g., two contacts over a two-year period give Fij = 1). The text would benefit from an explicit statement of how Fij is computed from Cij and the observation window.","section":"Section V-B"},{"comment":"The references to 'Figures VI-B and VI-B' in the text describing Figure 4 are placeholders that should be replaced with the actual figure panel labels.","section":"Section VI-B"},{"comment":"The description of the initial seed as a 'random user highly active in the first years of Mastodon' should specify how the random selection was performed; if the seed was manually chosen or convenience-selected, that should be stated.","section":"Section IV"},{"comment":"The paper claims that no prior work has analyzed Mastodon ego networks, but the text immediately acknowledges one prior study that used a graph-theoretic definition. The novelty statement should be sharpened to emphasize the distinction from the Dunbar ego-network model rather than claiming no prior ego-network study exists.","section":"Section III"},{"comment":"The 'Others2' group is defined as users active only after the acquisition, but this includes users who joined months or years later, not only those who migrated in the immediate post-acquisition wave. The term 'newcomers' in the title should be interpreted accordingly, or the analysis could be refined by cohort.","section":"Section VI-A"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about Table VII is well founded: the paper's own data show that most scaling ratios are not close to 3, and the outermost ratio is consistently much larger. This is a correctable problem if the authors reframe the central claim more precisely, but as written it affects the core message. The single-seed snowball design is also a serious external-validity limitation that should be acknowledged and preferably mitigated with additional seed analysis. The manuscript is a reasonable first step for the journal's scope, and I believe the authors can address the issues with a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honest take: the paper makes a real contribution by being the first to extract Dunbar-style ego networks from Mastodon interaction data, and the activity analysis around the Twitter acquisition is careful and useful. The central compatibility claim is only partially supported: the number of circles and the general layered structure do match the model, but the scaling-ratio evidence in Table VII is much weaker than the introduction suggests. Most consecutive-layer ratios are between 1.8 and 2.9 or between 4.4 and 5.9, not '~3'; only the middle ratios sit near 3. The paper acknowledges the outer-layer exception, but the '~3' gloss in Section I is misleading.\n\nWhat is genuinely new: prior Mastodon network studies used follower/followee graphs or a graph-theoretic ego network definition. This paper uses directed communication events, applies the standard ego-network pipeline (active threshold, six-month duration, Meanshift clustering), and finds 4-5 circles with increasing layer sizes. The post-acquisition activity analysis—separating Aficionados, Others1, Others2 and showing the Others2 group's directed communication patterns—is solid descriptive work.\n\nSoft spots, in order of importance. First, the sample is one snowball from a single seed on mastodon.social, capped at 2,000 egos, with no released data. That makes it hard to know whether the structure is typical of Mastodon or of that seed's community. Second, there are no error bars, no null model, and no sensitivity analysis around the filtering thresholds. The thresholds come from the authors' own prior papers; that's not a fatal circularity because the comparison target is the external Dunbar model, but it does mean the result is contingent on those choices. Third, restricting to post-acquisition users is reasonable but should be flagged more explicitly as a scope limitation.\n\nThese are fixable. The layered structure is probably real, but the 'canonical ~3 scaling' claim needs better statistics or softer wording. I'd send this to a serious referee; it's exactly the kind of descriptive evidence a good venue should evaluate, with the expectation of heavy revision.","headline":"First Dunbar-style ego network study on Mastodon with a solid activity analysis, but the scaling-ratio claim is overstated relative to the paper's own tables.","tokens_in":13776,"tokens_out":2762,"would_cite":false,"duration_ms":22724,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that Mastodon users' interaction networks form the same layered, Dunbar-style circles found on older platforms, with a scaling ratio of about three between layers.","keywords":["Mastodon","ego networks","Dunbar's number","social circles","Fediverse","snowball sampling","Meanshift clustering","decentralized online social networks"],"falsifier":"Re-run the extraction starting from several unrelated seeds on different instances and with a larger stopping threshold; if the modal circle count and the between-circle scaling ratio depart markedly from 4-5 circles and a ratio near $3$, the claim that Mastodon ego networks are canonically Dunbar-like would be falsified. A second falsifier is to apply the identical pipeline to pre-acquisition Aficionados: if their outer layers are just as underfilled as the newcomers', the 'young network' interpretation would lose its footing.","tokens_in":1574,"feed_emoji":"🐘","tokens_out":2354,"duration_ms":66037,"temperature":0.7,"pith_summary":"This paper sets out to show that even a young, decentralized social network like Mastodon organizes users' relationships into the same layered circles that have been found in offline social life and in older platforms like Twitter and Facebook. By analyzing direct interactions among roughly two thousand users who joined Mastodon after the 2022 Twitter acquisition, the authors find that most egos have four or five concentric circles of alters, with each outer circle about three times larger than the one inside it, in line with Dunbar's ego-network model. The smaller-than-canonical outer layers are interpreted as a sign of a network still forming, with weaker ties not yet fully built. If accepted, this makes Mastodon a workable open 'big data microscope' for studying human social behavior now that Twitter's API is no longer freely available.","feed_headline":"Mastodon users organize friends into Dunbar's 3x layers","feed_subtitle":"If the pattern holds, the open Fediverse could replace Twitter as a window into human social structure.","key_machinery":"The central object is the Dunbar ego network, a model where each individual (ego) has concentric circles of alters grouped by interaction frequency, with canonical sizes $1.5, 5, 15, 50, 150$ and a scaling ratio of about $3$. The paper constructs these networks from Mastodon interaction data, filters out alters contacted less than once per year or for less than six months, and then applies the Meanshift clustering algorithm (a non-parametric density-mode finder) to the annual contact frequencies so that the number of circles emerges without being forced. This pipeline yields the circle counts, layer sizes, and scaling ratios that the analysis compares against the canonical model.","core_discovery":"The central discovery is that ego networks on Mastodon—built from directed toots, replies, mentions, and boosts—exhibit the Dunbar-layered structure: a small number of social circles (most commonly 4 or 5) whose average alter counts increase by a factor close to $3$ from inner to outer circles, comparable to what has been measured on early Twitter and Facebook. The paper also finds that the active ego-network size is close to Dunbar's number (about 150 alters contacted at least once a year), while the full alter count is larger, and that external layers are comparatively underfilled, which it reads as evidence of young, still-developing networks among post-acquisition users.","pith_inferences":["Since the snowball sample grows from a single seed on mastodon.social, the reported structure may reflect that seed's local community; testing with multiple independent seeds on different instances would establish whether the 'young network' picture holds platform-wide.","The claim that weak ties need time to stabilise implies a testable prediction: for a fixed cohort, the sizes of the outer circles should grow relative to the inner ones across successive observation windows.","The paper excludes the pre-acquisition 'Aficionados' from the ego-network analysis; comparing their layers with the newcomers' would show whether the underfilled outer circles are specific to post-acquisition growth or common to all Mastodon users."],"forward_implications":["Mastodon can stand in for Twitter as an open platform for ego-network research, since its data come from a free public API.","Because the outer layers are underfilled, a replication in a few years should show circles growing toward the canonical 50 and 150 alters if the network matures.","The dominance of direct communication among post-acquisition users suggests that Mastodon is sustaining genuine social interaction rather than one-way broadcasting.","The successful extraction of Dunbar layers from Mastodon data opens the way to applying the same pipeline to other decentralized platforms, including Bluesky, which the authors name as a next step."],"supporting_citations":[{"why":"Supplies the standard ego-network extraction pipeline and the early Facebook/Twitter comparison that grounds the 'young network' reading.","marker":"[2]"},{"why":"Defines the canonical Dunbar ego-network model for online social networks, the benchmark for circle counts and scaling ratios.","marker":"[4]"},{"why":"Provides the Meanshift-based clustering pipeline and the Twitter ego-network analysis whose circle-count distribution is compared here.","marker":"[5]"},{"why":"The core reference establishing that online ego networks mirror offline ones, the premise for expecting Mastodon to match.","marker":"[8]"},{"why":"Source of the Dunbar number and the notion of active alters contacted at least once a year.","marker":"[14]"},{"why":"Establishes the discrete hierarchical organization of social group sizes with scaling ratio near 3 that the paper compares its ratios against.","marker":"[24]"},{"why":"Documents the post-acquisition migration from Twitter to Mastodon, supporting the focus on users who joined after the acquisition.","marker":"[12]"},{"why":"Defines the Meanshift algorithm used to extract social circles from interaction frequencies.","marker":"[6]"}],"fun_headline_variants":["Mastodon users' social circles follow Dunbar's 3x rule","Dunbar's layered structure confirmed on Mastodon egos","Mastodon's young networks still show Dunbar's 3x pattern","Mastodon mirrors Twitter's Dunbar layers for social study","Open Mastodon data reveals Dunbar's 3x ego network pattern"],"cache_read_input_tokens":15744,"weakest_assumption_plain":"The sample is one snowball connected component started from a single early, highly active user on mastodon.social, and the paper assumes this component is representative of Mastodon's overall newcomer population.","fun_headline_variants_meta":{"raw":{"variants":["Mastodon users' social circles follow Dunbar's 3x rule","Dunbar's layered structure confirmed on Mastodon egos","Mastodon's young networks still show Dunbar's 3x pattern","Mastodon mirrors Twitter's Dunbar layers for social study","Open Mastodon data reveals Dunbar's 3x ego network pattern"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000596,"raw_usage":{"total_tokens":2784,"prompt_tokens":937,"completion_tokens":1847,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":1753}},"tokens_in":553,"tokens_out":1847,"duration_ms":13220,"temperature":1.0,"reasoning_tokens":1753,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:37:29.704277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the extraction starting from several unrelated seeds on different instances and with a larger stopping threshold; if the modal circle count and the between-circle scaling ratio depart markedly from 4-5 circles and a ratio near $3$, the claim that Mastodon ego networks are canonically Dunbar-like would be falsified. A second falsifier is to apply the identical pipeline to pre-acquisition Aficionados: if their outer layers are just as underfilled as the newcomers', the 'young network' interpretation would lose its footing.","supporting_citations":[{"cited_title":"Computer Communications 76, 26–41 (2016)","cited_arxiv_id":null,"evidence_quote":"Supplies the standard ego-network extraction pipeline and the early Facebook/Twitter comparison that grounds the 'young network' reading."},{"cited_title":"Elsevier (2015)","cited_arxiv_id":null,"evidence_quote":"Defines the canonical Dunbar ego-network model for online social networks, the benchmark for circle counts and scaling ratios."},{"cited_title":"In: Companion Proceedings of the The Web Conference 2018","cited_arxiv_id":null,"evidence_quote":"Provides the Meanshift-based clustering pipeline and the Twitter ego-network analysis whose circle-count distribution is compared here."},{"cited_title":"Social Networks 43, 39–47 (2015)","cited_arxiv_id":null,"evidence_quote":"The core reference establishing that online ego networks mirror offline ones, the premise for expecting Mastodon to match."},{"cited_title":"Human nature 14(1), 53–72 (2003)","cited_arxiv_id":null,"evidence_quote":"Source of the Dunbar number and the notion of active alters contacted at least once a year."},{"cited_title":"Proceedings","cited_arxiv_id":null,"evidence_quote":"Establishes the discrete hierarchical organization of social group sizes with scaling ratio near 3 that the paper compares its ratios against."},{"cited_title":"In: Proceedings of the 2023 ACM on Internet Measurement Conference","cited_arxiv_id":null,"evidence_quote":"Documents the post-acquisition migration from Twitter to Mastodon, supporting the focus on users who joined after the acquisition."},{"cited_title":"IEEE Transactions on pattern analysis and machine intelligence 24(5), 603–619 (2002)","cited_arxiv_id":null,"evidence_quote":"Defines the Meanshift algorithm used to extract social circles from interaction frequencies."}],"review_version":1}