{"id":"7109d169-43e6-4e2a-a15e-a63f2da4ef0d","arxiv_id":"1908.10948","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A large-scale Twitter analysis shows geotagging is uneven across user groups, correlates with profile location reporting, and clusters in social networks, threatening assumptions behind geotagged-opinion research.","lead":"This paper measures how often Twitter users geotag their tweets, using 40 billion tweets from 20 million accounts. It finds that geotagging rates vary sharply by language, that users who list a location in their profile are more likely to geotag, and that people tend to befriend others with similar geotagging habits.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Homophily claim is confounded: Fig. 4 and Tables IX–X compare geotagged vs non-geotagged users without controlling for degree or activity, even though Fig. 3 shows geotag rates vary with follower counts.","rationale":"The reader's formal weakest assumption concerns sampling randomness and 3200-tweet truncation, while I focus on the degree/activity confound in the homophily analysis. These are related but distinct: both threaten the generalizability of specific magnitudes, but the homophily confound most directly undermines the paper's third central claim. The reader's rationale does mention the activity confound, so there is partial agreement, but it was not listed as the weakest assumption. I do not think this changes the overall verdict: the paper is a large descriptive study, the first two findings (language heterogeneity, profile-location correlation) are likely robust, and the homophily effect may survive correction, but its current evidence is not conclusive. A conditional verdict remains appropriate, and the proposed matching or permutation test is a feasible path toward strengthening or refuting the claim. I am not calling the result fraudulent; I am identifying a specific technical confound that the paper does not address.","tokens_in":12322,"tokens_out":3019,"duration_ms":33242,"concrete_test":"Recompute the Figure 4 distributions and the Table IX conditional probabilities within matched strata: bin users by follower count, followee count, and tweet count (or by number of alters with available data), and compare geotagged vs non-geotagged egos inside each bin. A sharper version is a degree-preserving null: randomly permute geotag status across egos while holding each ego's degree and activity fixed, then measure how often the observed KS statistic and odds ratios appear under this null. If the 37.41% vs 22.86% follower gap and the 6x conditional odds largely disappear within strata or under the permutation null, the homophily claim is an artifact of degree assortativity rather than evidence about geotagging preferences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The graph-level analysis in Section VI is the weakest support for the abstract's third claim of 'strong homophily effect.' In Figure 4 and Tables IX–X, geotagged and non-geotagged users are compared on the percentage of geotagged alters, but the analysis does not control for the ego's number of followers/followees, tweet volume, or activity level. Figure 3 itself shows that the probability of being geotagged rises rapidly with follower/followee count before declining, so geotagged users are not distributed uniformly over degree. Because social networks are degree-assortative, a geotagged ego with many alters will tend to have alters with many alters, and those alters will have a higher baseline geotag rate for reasons unrelated to geotagging preference homophily. The conditional-probability framing in Tables IX–X amplifies this: P(at least one alter is geotagged) increases mechanically with the number of alters, so comparing 'at least one alter geotagged' with 'no alter geotagged' without conditioning on degree conflates tie count with preference similarity. The paper notes in Section III.A that timelines are truncated at 3200 tweets, which further biases user-level geotag classification by activity, but the homophily confound is more directly load-bearing for the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a large-scale descriptive study of geotagging behavior on Twitter, based on roughly 40 billion tweets from about 20 million users. It makes three main claims: (1) different groups of users have very different geotagging rates, e.g., by language and client platform; (2) users who fill in a recognizable profile location are more likely to geotag their tweets; and (3) there is a strong homophily effect in geotagging preferences, so that geotagged users tend to be connected to geotagged users. The authors argue these findings challenge common assumptions behind opinion inference from geotagged tweets and behind location prediction systems. The paper is organized at three levels: tweet-level analysis, user-level analysis, and graph-level analysis, with the homophily claim supported by degree-based comparisons and conditional probabilities in Section VI.","tokens_in":12506,"tokens_out":3770,"duration_ms":38896,"significance":"If the findings hold, they are practically important: they imply that geotagged tweets are not a representative sample of public opinion and that classifiers trained on geotagged users may not transfer to non-geotagged users. The scale of the dataset is a genuine strength, and the descriptive tables are useful reference material. I also credit the authors for clearly separating tweet-level and user-level definitions of geotagging and for reporting the source-platform and language breakdowns in detail. The study is descriptive rather than model-based, so circularity is not a concern. However, two load-bearing issues prevent me from endorsing the strong conclusions as stated: the sampling and user-geotagging definitions are not as representative as claimed, and the homophily analysis does not control for degree or activity, even though the paper's own Figure 3 shows those variables are strongly associated with geotagging.","major_comments":[{"comment":"The 'strong homophily effect' claim is confounded by degree and activity. Figure 3 shows that the probability of being geotagged rises rapidly with the number of followers/followees before declining, so geotagged and non-geotagged users have systematically different degree distributions. The comparisons in Figure 4 and Tables IX-X do not condition on the ego's degree, tweet volume, or activity level. This matters for two reasons. First, if networks are degree-assortative, a geotagged ego with many alters tends to have alters with many alters, and those alters have a higher baseline geotag rate for reasons unrelated to geotagging-preference homophily. Second, the conditional-probability framing in Tables IX-X compares P(ego geotagged | at least one alter geotagged) with P(ego geotagged | no alter geotagged); the former increases mechanically with the number of alters, so the reported 'at least 6 times' ratios conflate tie count with preference similarity. I recommend re-analyzing the homophily claim with degree-stratified comparisons, controlling for the number of alters, or using a matched/regression design that holds degree and activity fixed.","section":"Section VI, Figure 4, Tables IX-X"},{"comment":"The sampling frame is not a uniform random sample of the Twitter population. The paper states in Section III.A, 'We consider this sampled streaming data representing a random sample of the Twitter population,' but the data actually consist of users who appear in Twitter's public sample stream and whose timelines and following lists are publicly accessible. Users with protected accounts, very low activity, or accounts not captured by the stream are excluded, and the stream itself is not a uniform random sample of users. In addition, the user-level definition of geotagging relies on the most recent 3,200 tweets per user, so heavy users' older geotags are systematically missed. These issues directly affect the headline user-level rates reported in Section III.B (e.g., 24.38% geotagged users) and the language/source breakdowns in Tables VI and VII, because activity level is correlated with tweet volume. Please either provide explicit evidence for the representativeness assumption, restrict the claims to the sampled population, or report robustness checks stratified by tweet volume and account age.","section":"Section III.A and Section III.B, Table I"},{"comment":"The Kolmogorov-Smirnov tests in Section VI are run on millions of observations, so p<0.001 is expected and not informative. The reported KS statistics (0.255 to 0.431) are moderate, but the text interprets them as showing 'significant difference' without discussing effect sizes or practical significance. I recommend reporting standardized mean differences, overlapping coefficients, or other effect-size measures, and acknowledging that with this sample size any small distributional shift will be statistically significant.","section":"Section VI, KS tests"}],"minor_comments":[{"comment":"The phrase 'may affects the generability' should be corrected to 'may affect the generalizability.'","section":"Abstract"},{"comment":"The text spells the bot platform as 'twitterbot.net' while Table II lists 'twittbot.net'; please make the names consistent. The footnote URL in Section III.A is also malformed ('en9/docs').","section":"Table II and Section III.A"},{"comment":"The figure has dual axes but the axes are not labeled; please label the left and right axes explicitly so the reader can interpret the bar chart and the trend lines.","section":"Figure 1"},{"comment":"Table V reports coordinates-tagged percentages among geotagged tweets only, because non-geotagged tweets lack country labels. The text should state this more prominently so readers do not interpret the table as country-level geotagging rates.","section":"Section IV.B, Table V"},{"comment":"The analysis restricts to users with at least five followers/followees, but the threshold is not justified. Please report sensitivity of the homophily results to this threshold.","section":"Section VI"},{"comment":"The concluding sentence says the study is based on '20 million randomly sampled Twitter users,' but given the caveats in Section III.A, 'randomly sampled' is too strong; please qualify the description.","section":"Section VII"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a valuable descriptive study of geotagging behavior on an impressive scale, but the headline homophily result is supported by an analysis that doesn't control for degree or activity, and that weakens the third claim considerably.\n\nWhat's actually new: the paper documents, on 40B tweets from 20M users, that geotagging rates vary sharply by language (Korean under 3%, Indonesian over 40%), that users who self-report a recognizable profile location are much more likely to geotag, and that there is a positive association between a user's geotag status and the geotag rate of their followers/followees. The profile-location correlation is a genuine addition beyond Sloan and Morgan, and the scale makes the tweet- and user-level descriptive tables robust in the ways that matter.\n\nWhere it gets soft: the sampling frame. The authors call the streaming API sample 'random,' but it's a convenience sample of users active during the collection window, and the 3,200-tweet cap means 'ever geotagged' is really 'geotagged within the last N tweets.' That biases user-level rates and likely inflates the correlation with activity.\n\nThe bigger issue is the homophily analysis. Figure 3 shows that geotagging probability rises sharply with follower/followee count, then declines. Geotagged users are therefore not distributed uniformly over degree. Since networks are degree-assortative, a geotagged ego with many alters tends to have alters with many alters, and those alters have a higher baseline geotag rate for reasons unrelated to preference similarity. The conditional probabilities in Tables IX-X make it worse: P(ego geotagged | at least one alter geotagged) mechanically increases with alter count, so comparing it to P(ego geotagged | no alter geotagged) without conditioning on degree conflates tie count with homophily. The paper needs to stratify or regress on degree and activity before the 'strong homophily effect' claim is credible.\n\nWhere the paper holds up: the descriptive core—language and source differences, the profile-location link, the overall geotag prevalence—is solid and likely robust to the sampling caveats. The authors also clearly flag the 3,200 cap as a limitation, which is honest. No code or data is released, which limits reproducibility but doesn't undermine the descriptive tables.\n\nWho should read it: anyone building or using text- or graph-based geolocation predictors, and anyone who uses geotagged tweets as a proxy for local opinion. It's a useful caution even if the homophily magnitude is provisional.\n\nRecommendation: send it to peer review. The descriptive findings deserve referee time, and the homophily section needs reanalysis, not rejection. I'd ask for degree/activity controls and a more careful treatment of the sampling frame before accepting.","headline":"Large descriptive study with a valuable cautionary message, but the headline homophily claim is confounded by degree and activity and needs reanalysis.","tokens_in":13086,"tokens_out":2362,"would_cite":true,"duration_ms":20696,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Geotagged tweets are not a random sample of public opinion: language alone predicts geotagging from under 3% (Korean) to over 40% (Indonesian).","keywords":["geotagging","Twitter","user behavior","empirical study","homophily","location prediction","language differences","social media"],"falsifier":"Use Twitter's full-archive search or an account's downloadable data to inspect the complete tweet history, including tweets older than the platform's 3,200-tweet limit, for a random sample of accounts the paper classifies as non-geotagged. If a meaningful fraction of those accounts have geotagged earlier, then the ever-geotagged rate and the friend-homophily conditional probabilities are truncated-history artifacts rather than true behavioral patterns.","tokens_in":12066,"feed_emoji":"📍","tokens_out":8191,"duration_ms":70674,"temperature":0.7,"pith_summary":"This paper tests a quiet assumption behind much social-media research: that tweets carrying a location tag are a fair sample of what people are saying. Using 41 billion tweets from 20 million Twitter users, it shows the assumption fails. Geotagging is sharply uneven across groups—fewer than 3% of Korean-speaking users ever geotag, while more than 40% of Indonesian-speaking users do—and it tracks other disclosure choices: people who write a recognizable home location in their profile are far more likely to share real-time coordinates. The same preference also runs in social ties, since users connected to a geotagging friend are about six times more likely to geotag than users with no such friend. If this is right, opinion estimates and location-prediction systems built only on geotagged users carry a built-in bias toward people who choose to be visible.","feed_headline":"Under 3% of Korean users geotag; over 40% of Indonesian users do","feed_subtitle":"The bias matters for opinion studies and location classifiers trained on geotagged users.","key_machinery":"The central object is a user-level geotagging label: after collecting each user's timeline, the authors call a user geotagged if any tweet has a place tag or precise coordinates, and coordinates-tagged if any tweet has precise coordinates. Around that label they build three partitions—major tweet source, major tweet language, and whether a free-text profile location resolves to a place through a public gazetteer—and compare geotagging rates across partitions. The argument is carried by the combination of these partitions with standard statistical comparisons (chi-square tests for rate differences and Kolmogorov-Smirnov tests for the distribution of geotagged friends), plus the ego-alter conditional probabilities that express the homophily effect as a sixfold increase.","core_discovery":"On its own terms, the paper claims that geotagging behavior on Twitter is a systematic user-level trait rather than a random property of tweets. Over the collected 41.27 billion tweets, only 2.31% carry a place tag or coordinates, but when aggregated per user, 24.38% of the 19.98 million users have geotagged at least once and 12.93% have shared precise coordinates. The rate varies by platform (iPhone users geotag more than Android users), by country (Indonesia and Turkey show exceptionally high coordinates-tagging), and sharply by language: 2.94% of Korean-speaking users are geotagged versus 42.47% of Indonesian-speaking users. Users who report a recognizable location in their profile are also much more likely to be geotagged (33.03%) than users with an empty profile field (15.91%). Finally, geotagging is homophilous: geotagged users' followees and followers are geotagged at roughly 12 to 15 percentage points higher than the corresponding ties of non-geotagged users, and the conditional probability that a user geotags jumps from about 4% to roughly 26-31% once at least one connected user geotags. The authors take these patterns as evidence against the common assumption that geotagged and non-geotagged users are exchangeable.","pith_inferences":["The person-level correlation between profile disclosure and geotagging points to a stable privacy-sensitivity trait; one could test this by checking whether users who disable geotagging also withhold profile details and engage less with location-based features.","The same selection mechanism likely distorts other volunteered geographic information beyond Twitter, such as check-in services or photo geotags, so reweighting by user-level propensity could improve cross-platform location studies.","The observed homophily could be confounded by spatial homophily: people who live near each other share language and culture and both geotag; a direct test would compare geotagging similarity between friends who live far apart, where social influence is the only channel."],"forward_implications":["Regional opinion studies that use geotagged tweets as a proxy for local sentiment should reweight text by language- and country-specific geotagging likelihood instead of treating all geotags as equal draws.","Location-prediction classifiers trained on geotagged users inherit a selection bias, because users who geotag are also disproportionately the users who fill in their profile location, making the training labels non-representative of the non-geotagged users being predicted.","Graph-based location inference becomes harder than assumed: if non-geotagged users cluster with non-geotagged users, then friendship ties carry less location signal for exactly the users who most need location inference.","Cross-country comparisons of tweet volume on a topic can be misleading, since the same real-world concern can produce far more geotagged discussion in Turkey or Indonesia than in Japan or Korea purely because of geotagging propensity.","The strong friend correlation suggests geotagging behavior may spread through social influence, so studies of location-based social behavior should account for network autocorrelation."],"supporting_citations":[{"why":"Provides the closest prior demographic analysis of geotagging on a 1% streaming feed, which this paper extends with per-user timelines and following ties.","marker":"[27]"},{"why":"Supplies the earlier estimate of geotagged tweet volume that this paper's 2.31% figure updates.","marker":"[9]"},{"why":"Represents the neural classifiers trained on geotagged tweets whose generalizability the findings question.","marker":"[10]"},{"why":"Defines the homophily concept that the friend-similarity analysis operationalizes.","marker":"[11]"},{"why":"Documents geographical homophily in movement data and motivates the graph-based location prediction that the findings call into question.","marker":"[7]"},{"why":"Identifies the self-reported profile location as the most important feature for user geolocation, giving the profile-location correlation its practical stakes.","marker":"[30]"}],"fun_headline_variants":["Geotagging is a habit: only 2% of tweets, but 24% of users tag","Twitter geotagging: language, profile, and friendship predict it","Geotagging bias: Indonesian 42% vs Korean 3%, and it's contagious","Geotagging on Twitter is systematic, not random: friends matter","24% of users geotag, but only 2% of tweets carry a location"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that the 20 million users drawn from Twitter's streaming API are a random sample of all Twitter users, and that the most recent 3,200 tweets per user reveal every user's true geotagging behavior; if active users or truncated histories are unrepresentative, every rate and the homophily effect could be off.","fun_headline_variants_meta":{"raw":{"variants":["Geotagging is a habit: only 2% of tweets, but 24% of users tag","Twitter geotagging: language, profile, and friendship predict it","Geotagging bias: Indonesian 42% vs Korean 3%, and it's contagious","Geotagging on Twitter is systematic, not random: friends matter","24% of users geotag, but only 2% of tweets carry a location"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000429,"raw_usage":{"total_tokens":2240,"prompt_tokens":1037,"completion_tokens":1203,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1090}},"tokens_in":653,"tokens_out":1203,"duration_ms":11705,"temperature":1.0,"reasoning_tokens":1090,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:28:48.943946+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use Twitter's full-archive search or an account's downloadable data to inspect the complete tweet history, including tweets older than the platform's 3,200-tweet limit, for a random sample of accounts the paper classifies as non-geotagged. If a meaningful fraction of those accounts have geotagged earlier, then the ever-geotagged rate and the friend-homophily conditional probabilities are truncated-history artifacts rather than true behavioral patterns.","supporting_citations":[{"cited_title":"Who tweets with their location? understanding the relationship between demographic characteristics and the use of geoservices and geotagging on twitter,","cited_arxiv_id":null,"evidence_quote":"Provides the closest prior demographic analysis of geotagging on a 1% streaming feed, which this paper extends with per-user timelines and following ties."},{"cited_title":"Where in the world are you? geolocation and language identiﬁcation in twitter,","cited_arxiv_id":null,"evidence_quote":"Supplies the earlier estimate of geotagged tweet volume that this paper's 2.31% figure updates."},{"cited_title":"On predicting geolocation of tweets using convolutional neural networks,","cited_arxiv_id":null,"evidence_quote":"Represents the neural classifiers trained on geotagged tweets whose generalizability the findings question."},{"cited_title":"Birds of a feather: Homophily in social networks,","cited_arxiv_id":null,"evidence_quote":"Defines the homophily concept that the friend-similarity analysis operationalizes."},{"cited_title":"Friendship and mobility: user movement in location-based social networks,","cited_arxiv_id":null,"evidence_quote":"Documents geographical homophily in movement data and motivates the graph-based location prediction that the findings call into question."},{"cited_title":"Text-based twitter user geolocation prediction,","cited_arxiv_id":null,"evidence_quote":"Identifies the self-reported profile location as the most important feature for user geolocation, giving the profile-location correlation its practical stakes."}],"review_version":1}