{"id":"42ad9e78-7dec-480f-a883-8fc085ef7d41","arxiv_id":"2507.12372","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Web-browsing LLMs can retrieve X profile content and infer demographics with above-chance accuracy in some cases, but the study's evidence is partly confounded by training-data memorization and a heavily reduced synthetic sample.","lead":"This paper tests whether web-browsing AI chatbots can look up X (Twitter) profiles and guess a user's age, gender, income, and politics from the username alone. The authors find the models often can, but the evidence is weaker than the headline accuracy numbers suggest, especially because the AI may have already learned about many accounts during training.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey-dataset result may reflect training-data memorization, not live web browsing; without a no-browsing control, the central capability claim is unsupported.","rationale":"The reader's weakest_assumption identifies exactly the concern that is most load-bearing: the survey dataset is contaminated by potential training-data memorization, and no control isolates the web-browsing mechanism. This is not a minor methodological footnote; it determines whether the paper demonstrates a new capability or merely re-demonstrates that LLMs retain social media content they have already seen. The synthetic dataset, while not subject to this confound, is too small and too heavily reduced (27 accounts) to carry the headline claim alone. A no-browsing control is straightforward and would settle the issue. Since the reader already flags this as the basis for a CONDITIONAL verdict, my independent stress-test does not change the verdict; it reinforces it.","tokens_in":128,"tokens_out":2209,"duration_ms":46279,"concrete_test":"Run the identical demographic inference prompts on the same 1,384 survey handles under two conditions: (1) default web-browsing enabled, and (2) web access disabled (e.g., a non-browsing model deployment). Compare classification accuracy per attribute and model. If the no-browsing accuracy is statistically indistinguishable from the browsing accuracy, the observed performance is attributable to pretraining memorization rather than live retrieval; if browsing accuracy is significantly higher, the confound is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest evidence for web-browsing-enabled demographic inference comes from the 1,384-account survey dataset, on which GPT-4o and GPT-o3 exceed 50% accuracy on most attributes. However, Section 3.1 explicitly states that these accounts 'have already been used to train Grok as well as, most likely, other models such as ChatGPT.' Because the prompts ask the model to infer demographics 'based on tweets posted by handle/link,' the model could answer from parametric memory of those tweets, without performing any live retrieval. The synthetic accounts avoid this confound because they were newly created, but only 27 of 48 remained usable and the paper labels this a lower bound; the high-accuracy headline results therefore rest on the confounded dataset. The paper reports no control condition with web browsing disabled, no retrieval logs, and no verification that the models actually fetched current profile content rather than recalling training data. Without such a control, the central claim that web browsing enables this profiling capability is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper investigates whether web-browsing large language models (GPT-4o, GPT-o3, and Llama-3-8B-Web) can, given only an X/Twitter handle, retrieve profile content and infer a user's age, gender, socioeconomic status, and political orientation. The authors use two datasets: 48 synthetic accounts created for the study (27 remain after suspensions) and 1,384 accounts from a 2018 survey with self-reported demographics. They report above-chance accuracy on the survey dataset, explore which profile signals the models rely on using the synthetic accounts, and discuss privacy, security, and policy implications. The central claim is that web-browsing LLMs enable zero-data social media profiling with reasonable accuracy.","tokens_in":23529,"tokens_out":4771,"duration_ms":56733,"significance":"If the capability claim is established, this is a novel and important result for computational social science and for privacy research: it would show that an LLM with live web access can profile individuals from a username alone, with no API access. The paper has clear strengths: a controlled synthetic-account design with pre-specified demographic profiles, careful attention to research ethics (private accounts, no interaction with users, anonymized survey data), honest reporting of limitations such as account suspension and lower-bound interpretation, and a useful mechanism analysis. The synthetic-account portion is a genuinely interesting probe of how models use profile pictures, bios, and following lists. However, the quantitative headline results rest primarily on the survey dataset, which is acknowledged to overlap with the models' training data, and no control separates live browsing from parametric memory. The current evidence is therefore insufficient to support the paper's central claim at the stated strength.","major_comments":[{"comment":"The survey-dataset result is confounded by training-data memorization, and this is load-bearing for the paper's central claim. Section 3.1 explicitly states that the 2018 accounts 'have already been used to train Grok as well as, most likely, other models such as ChatGPT.' Because the prompts in Table 1 ask the model to answer 'Based on tweets posted by handle/link,' the model can answer from parametric memory of those tweets without performing any live retrieval. The paper reports no control condition with web browsing disabled, no retrieval logs, and no independent verification that the models actually fetched current profile content. Consequently, the high accuracies in Figure 2 and Tables 7–8 do not establish that web browsing enables the profiling capability; they may reflect memorized training data. This is the central unsupported claim of the paper.","section":"§3.1, §4.2"},{"comment":"The synthetic dataset, which avoids the training-data confound, is too small and too attrition-prone to support the mechanism claims. Of the 48 accounts, 21 were suspended, leaving 27, and all 16 accounts in the no-bio condition were suspended. This means the controlled manipulation of bios cannot be evaluated at all, and the comparisons in Figure 1b–1d (profile picture, following, suspended vs. non-suspended) rest on subgroups of the 27 accounts with no statistical testing. The paper repeatedly notes the small sample and says results are exploratory, but Section 4.1 presents these comparisons as findings (e.g., 'the presence of a profile picture improves gender prediction accuracy across all three models'). As presented, the evidence cannot distinguish true signals from noise, and the loss of the no-bio condition removes a key part of the intended design.","section":"§4.1, §3.1"},{"comment":"The choice of the Chatbot-Handle prompting combination is post hoc and likely inflates reported accuracy. The paper states that no single prompt-identifier combination consistently outperformed others, then selects Chatbot-Handle for the main results. However, Appendix C shows that other combinations often perform better: for GPT-4o on the survey dataset, System-Handle gives age accuracy 0.83 vs. 0.72 for Chatbot-Handle, and User-Handle gives socioeconomic accuracy 0.90 vs. 0.88. For GPT-o3 on the synthetic dataset, System-Link gives political orientation 0.62 vs. 0.52 for Chatbot-Handle. Reporting the best-performing combination without pre-registration or multiple-comparison correction makes the headline accuracy numbers optimistic. The authors should either justify the selection criterion on a priori grounds or report all six combinations as primary analyses.","section":"§4.1, §4.2, Appendix C"},{"comment":"Accuracy comparisons to 'random chance' and significance claims are not statistically substantiated. After category collapsing, chance levels are 50% for gender, 25% for age (4 classes), and 33% for socioeconomic status and political orientation (3 classes). GPT-o3's political accuracy of 0.42 is only 9 points above chance, yet Figure 2a describes accuracies as 'substantially above random chance' without confidence intervals or hypothesis tests. The text also uses 'significantly' for model comparisons and for the conservative-user decline (Figure 2c, 2d), but no significance tests, confidence intervals, or effect sizes are reported for any accuracy value. This is a load-bearing issue because the paper's comparative claims (e.g., which model is best, which groups are biased) are presented as findings without inferential support.","section":"§4.2, §3.3"}],"minor_comments":[{"comment":"The figure caption is inconsistent with the panel labels: the caption lists (c) as 'accounts with and without bios' while the panel itself is labeled 'Following Vs. Non-Following,' and panels (e) and (f) are not described in the caption. Please align the caption with the actual panel content.","section":"Figure 1 caption"},{"comment":"There are several typos and inconsistencies: 'perfromance' in Section 2; 'and and' in Section 4.2; 'the different is not significant' should be 'the difference is not significant'; 'GPT-3.5 (o3)' should be 'GPT-o3'; 'Amozon' in the Appendix tweets; and 'underrate' in a tweet appears truncated.","section":"Throughout"},{"comment":"The paragraph on IRB approval argues that the survey data are publicly available and that X's terms permit use for model training, but the models are not 'trained' by the research; the paper uses them in inference. This does not affect the technical results, but the framing could be tightened.","section":"Section 3.1"},{"comment":"The claim that suspended accounts have 'only the profile picture and screen name visible' is used to infer that models rely on names and pictures, but this inference depends on the assumption that the models do not have cached or parametrically memorized content from before suspension. Given the training-data concern raised elsewhere, this assumption should be acknowledged.","section":"Section 4.1, 'Suspended Accounts'"},{"comment":"The system prompt in Figure 3 contains a grammatical error ('You live is the USA') and defines age ranges that are inconsistent with the categories in Table 2 (e.g., 'below 14 years old' vs. 'Child (between 13 and 18)'). Please harmonize these definitions.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is on a timely and potentially high-impact topic, but the central quantitative claim currently rests on a dataset that is acknowledged to overlap with model training data, and the only unconfounded dataset is too small. I believe this can be addressed with additional control experiments (e.g., browsing disabled vs. enabled, fresh accounts created after the models' knowledge cutoff, retrieval-log verification) and a more cautious interpretation of the survey results. The scope of the manuscript is appropriate for a CS/CL venue, but the current evidence is not yet strong enough for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know before spending time on this: the paper is the first I've seen that tests web-browsing LLMs retrieving live X profiles from usernames alone and inferring demographics. That's a real gap in the LLM-as-profiler literature, and the authors deserve credit for building a controlled synthetic dataset to avoid training-data leakage. But the two datasets pull in opposite directions: the synthetic set is too small after suspensions (27 of 48, with the entire no-bio condition gone), and the survey set is almost certainly contaminated by memorization—the authors themselves note these accounts were likely used to train ChatGPT and Grok. Without a no-browsing control or retrieval logs, the high accuracies on 1,384 accounts cannot be attributed to live browsing.\n\nWhat's new and worth keeping: the controlled experiment with dummy accounts, the prompt/identifier comparisons, and the exploratory analysis of how models use profile pictures, bios, and follow lists. The observation that models attempt inference even on suspended accounts, relying on names, is a nice mechanistic clue. The paper is also unusually candid about its own limitations, which I respect.\n\nThe soft spots are real. First, the synthetic dataset loses all accounts in the no-bio condition, so one of the three experimental cells is missing; the remaining 27 accounts give wide intervals and no significant effects. The authors call this exploratory, which is fair, but it undercuts the bias claims. Second, the survey dataset is the main quantitative evidence, and it's confounded. The prompts say 'based on tweets posted by handle/link,' but the model could answer from parametric memory. No control with browsing disabled, no retrieval logs, no verification of what was actually fetched. Third, the paper reports accuracy without confidence intervals or a chance baseline that accounts for class imbalance; on a 4-class age task, 50% may not be as impressive as it sounds. These are fixable, but they need to be fixed before the central claim is credible.\n\nWho it's for: people working on LLM agents, web browsing, and computational social science methods. It would be a useful reading-group paper because the methodological debate is instructive. I'd cite the synthetic-dataset design, but not the survey accuracy numbers, until the control condition is added.\n\nRecommendation: send it to peer review, but with a clear request for a no-browsing control, retrieval logs, and baseline comparisons. The topic is important enough that referee time is justified even though the current evidence is conditional.","headline":"Plausible new capability, but the headline accuracy rests on a memorization confound; the synthetic data are too thinned out to carry the load.","tokens_in":24076,"tokens_out":2352,"would_cite":true,"duration_ms":27132,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Web-browsing LLMs can retrieve X profiles and infer user demographics from usernames alone, with accuracy that varies by model and attribute.","keywords":["large language models","web browsing","social media profiling","demographic inference","X (Twitter) accounts","privacy risks","computational social science","user modeling"],"falsifier":"Take a fresh cohort of X accounts created after every relevant model's training cutoff, with known demographics and fully public content; prompt the same models by username before and after making the accounts crawlable. If accuracy collapses when the accounts cannot be retrieved, the capability is genuine browsing; if accuracy persists on blocked or nonexistent handles, the dominant mechanism is memorized or hallucinated content.","tokens_in":23163,"feed_emoji":"🕵️","tokens_out":5021,"duration_ms":57721,"temperature":0.7,"pith_summary":"This paper asks whether web-browsing large language models can do what older classifiers could not: start from a bare X (formerly Twitter) handle, retrieve that person's live profile content, and infer their age, gender, socioeconomic status, and political orientation. The authors report that GPT-4o and GPT-o3 do this with accuracy well above chance on a 1,384-user survey dataset, and that even a smaller open model attempts the retrieval. A controlled experiment with 27 synthetic accounts shows the models draw on profile pictures, bios, and follow lists, and that they fall back on names even for suspended accounts, which can introduce gender and political bias. The paper positions this as a new zero-data profiling capability that could sustain computational social science after API restrictions or enable mass surveillance and targeting, and it recommends restricting the capability in public products.","feed_headline":"LLMs can browse X profiles and infer user demographics","feed_subtitle":"Given only a username, GPT-4o and GPT-o3 retrieve live tweets and predict age, gender, class, and politics.","key_machinery":"The operative mechanism is the web-browsing retrieval loop: the model receives a prompt containing only the account handle or profile link, internally fetches the live X page, and conditions its demographic answer on the fetched content including posts, bio, profile picture, replies, retweets, and follow lists. The paper's best-performing configuration is the 'Chatbot-Handle' prompt, which asks the model to complete the sentence \"Based on tweets posted by handle, given the options [...], I think the [attribute] of this user is:\" and which matched or outperformed other prompt and identifier combinations across models and datasets. The controlled synthetic design is what lets the authors attribute inferences to specific profile components: one-third of accounts had picture plus bio, one-third only bio, one-third only picture, and half followed 25 politically aligned accounts, so accuracy differences across those conditions reveal which cues the models actually use.","core_discovery":"The central claim is that browsing-enabled LLMs constitute a working profiler: given only a username, they access the account's posts, bio, replies, retweets, and profile picture, and predict demographic attributes with moderate-to-high accuracy. On the 1,384-user survey dataset, GPT-4o reaches 72–88 percent accuracy on age, gender, and socioeconomic status, and GPT-o3 reaches 90 percent on age; political orientation is lower but still mostly above chance. On the 27 usable synthetic accounts, accuracy is lower and is treated by the authors as a lower bound, with gender the most reliably inferred attribute and age the hardest. The paper also claims that the mechanism is visible in the synthetic experiments: models rely more on names and profile pictures than on textual history, which produces bias against minimal-activity accounts and, for GPT-4o, against conservative users.","pith_inferences":["The paper's memorization caveat suggests a strong test the authors did not run: if accounts are suspended or blocked between prompt and inference and accuracy persists, the survey results would indicate training-data recall rather than live browsing.","The same mechanism should transfer to other platforms with public profiles, which would widen both the research upside and the surveillance surface; a systematic multi-platform benchmark is the natural next step.","The reliance on names and profile pictures for gender and ideology implies that pseudonymous handles or neutral avatars could degrade inference accuracy, a testable privacy-preserving design.","If political-orientation accuracy stays low on real accounts, targeted political advertising from usernames may be less effective than feared, whereas gender and age targeting would remain the sharper risk."],"forward_implications":["Demographic inference from a username is now a live, low-cost operation, so the privacy risk is no longer hypothetical: anyone with API access can profile users at scale without an API feed.","Researchers in the post-API era can use browsing LLMs to reconstruct user-level attributes for studies that would otherwise lack data, provided ethical review and consent are handled.","Because models infer gender and ideology from names and pictures even when the account is suspended, profiles with minimal activity inherit systematic misclassification bias.","The accuracy gap between synthetic and survey accounts implies that real, organically maintained accounts give the models richer signals, so performance in natural deployments should be expected to be higher, not lower."],"supporting_citations":[{"why":"Supplies the survey dataset of 1,384 international X users with self-reported demographics against which the main accuracy results are measured.","marker":"[25]"},{"why":"Provides the user/system/chatbot prompt taxonomy and the profile-generation prompt adapted to create the synthetic personas.","marker":"[9]"},{"why":"Prior work showing that a vision-capable LLM can infer demographics from partisan tweets; this is the capability the paper extends to live retrieval from usernames.","marker":"[24]"}],"fun_headline_variants":["Browsing LLMs expose user demographics from X usernames","LLM profilers: just a username reveals age, gender, politics","Web-browsing LLMs infer demographics from social media handles","GPT-4o profiles X users: 72-88% accuracy on demographics","Browsing LLMs read X profiles to predict user traits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey-dataset results count as evidence for live web browsing only if the models are not instead recalling profiles they have already memorized from training data; the paper itself notes the 2018 accounts were likely used to train Grok and possibly ChatGPT, and no control condition separates browsing from memory.","fun_headline_variants_meta":{"raw":{"variants":["Browsing LLMs expose user demographics from X usernames","LLM profilers: just a username reveals age, gender, politics","Web-browsing LLMs infer demographics from social media handles","GPT-4o profiles X users: 72-88% accuracy on demographics","Browsing LLMs read X profiles to predict user traits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000182,"raw_usage":{"total_tokens":1304,"prompt_tokens":930,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":546,"completion_tokens_details":{"reasoning_tokens":281}},"tokens_in":546,"tokens_out":374,"duration_ms":4375,"temperature":1.0,"reasoning_tokens":281,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:47:16.816934+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh cohort of X accounts created after every relevant model's training cutoff, with known demographics and fully public content; prompt the same models by username before and after making the accounts crawlable. If accuracy collapses when the accounts cannot be retrieved, the capability is genuine browsing; if accuracy persists on blocked or nonexistent handles, the dominant mechanism is memorized or hallucinated content.","supporting_citations":[{"cited_title":"Cognitive reflection correlates with behavior on twitter","cited_arxiv_id":null,"evidence_quote":"Supplies the survey dataset of 1,384 international X users with self-reported demographics against which the main accuracy results are measured."},{"cited_title":"Gpt-4v (ision) as a social media analysis engine","cited_arxiv_id":null,"evidence_quote":"Prior work showing that a vision-capable LLM can infer demographics from partisan tweets; this is the capability the paper extends to live retrieval from usernames."}],"review_version":1}