{"id":"252693c3-692f-4698-8ccb-b17e5915b837","arxiv_id":"2411.10609","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"New anonymized datasets for 26 verified information operations include platform-confirmed IO posts and 13M+ timeline posts from control accounts selected by shared hashtags.","lead":"This paper releases new labeled datasets on 26 state-backed information operations, pairing verified inauthentic posts with more than 13 million control posts from organic accounts. It provides a shared benchmark resource for researchers building and testing influence-operation detection tools.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Control accounts are not observationally comparable to IO accounts (full timelines vs 100-post daily crops), creating a likely trivial confound for any account-level detection benchmark; the paper acknowledges but does not quantify it.","rationale":"Read in good faith, the paper's minimal claim is that the datasets exist and contain the stated numbers of IO and control posts. That claim appears true and internally consistent; the control-account total roughly sums to 303k. The paper also discloses several limitations, which is commendable. However, the stronger implied claim is that the datasets enable benchmarking of IO detection algorithms. That requires the control group to be a usable negative class. The most load-bearing weak point is not the hashtag-proxy selection bias (which the reader identified) but the observation-window asymmetry: IO accounts are observed over their full lifetimes, control accounts only on selected days with a 100-post cap and no posts after the inclusion date. This is not a subtle labeling issue; it is a structural difference in what data exists for each class. Any account-level feature that depends on activity volume or temporal extent will be confounded. The paper mentions this in the Discussion but does not quantify it or provide adjusted versions. A concrete baseline experiment can settle whether the artifact is severe; until then, the benchmarking claim should be treated as conditional. This does not change the reader's CONDITIONAL verdict but sharpens the condition.","tokens_in":11115,"tokens_out":5101,"duration_ms":52400,"concrete_test":"Download the released Zenodo data. For each campaign, compute per-account features from the raw data: total posts, number of distinct active dates, and account age (last observed date minus account creation date). Fit a logistic regression on these three features alone to predict the is_control label, using cross-validation. If the AUC is above approximately 0.9 in most campaigns, the temporal/volume artifacts almost fully separate IO from control accounts, demonstrating that any account-level detection benchmark built on this dataset is confounded and needs matching or reweighting before use.","verdict_should_be":"UNCHANGED","load_bearing_attack":"To support the central claim that researchers can benchmark IO detection algorithms, control accounts must be a valid negative class with observation windows comparable to IO accounts. The collection procedure in 'Control Data Collection' violates this: IO accounts contribute their full platform-verified timelines, while control accounts contribute only posts on dates when they used an IO hashtag, capped at 100 posts per day, with all later posts excluded. The Discussion states: 'control account timelines are cropped at 100 posts and exclude posts that occurred after the date on which they met the inclusion criteria.' Consequently, per-account features such as total post volume, number of active days, and observed account age are systematically higher for IO accounts. A detector trained on the released data could achieve high accuracy by exploiting these collection artifacts rather than by learning IO behavior. The authors acknowledge the temporal misalignment and suggest users 'selectively choose IO and control accounts with matching temporal activity,' but the released datasets are not adjusted, and no quantification of the confound is provided. Because the central contribution is benchmarking value, this unquantified artifact is load-bearing: if the artifact alone separates the classes, detection results on these datasets cannot be interpreted as measuring IO detection skill.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This dataset paper introduces and documents a new collection of labeled social-media datasets covering 26 state-sponsored information operations (IOs) attributed to 16 state actors. For each campaign, the release contains platform-verified IO posts (full account timelines) and a control set of posts by accounts that used the same hashtags as the IO accounts on the same dates. The control data reportedly comprise over 13 million posts from 303k accounts. The data are anonymized, merged into a common schema with an `is_control` flag, and released on Zenodo with a DOI. The paper describes the collection pipeline, provides descriptive statistics in Table 1, reports coverage statistics in Fig. 2, and discusses limitations including temporal misalignment and hashtag-proxy quality. The stated contribution is that these datasets enable researchers to characterize IO tactics and benchmark IO detection algorithms without needing to re-hydrate posts through now-inaccessible APIs.","tokens_in":11289,"tokens_out":7486,"duration_ms":73968,"significance":"If the control data are valid, the release is a useful community resource: it spans more campaigns and countries than most previous control sets, avoids the need for API access, and includes anonymized post content rather than only IDs. Strengths include the breadth of the release, the explicit coverage statistics, the careful acknowledgment of several limitations, and the inclusion of a data DOI and ethical documentation. The paper is also internally consistent in its reported aggregate totals: the sums in Table 1 approximate the abstract's 13M+ control posts. However, the central benchmarking claim depends on the control accounts being a usable negative class. The structural asymmetry between full IO timelines and cropped control timelines, plus the dependence on hashtag quality, means the datasets as released can support descriptive comparative studies but need additional validation or qualification before they can support reliable detection benchmarks.","major_comments":[{"comment":"The main load-bearing issue for the benchmarking claim is the structural asymmetry between IO and control timelines. In 'Control Data Collection', IO accounts contribute their full platform-verified timelines, whereas control accounts contribute only up to 100 messages per day on dates when they used an IO hashtag, and the Discussion states that 'control account timelines are cropped at 100 posts and exclude posts that occurred after the date on which they met the inclusion criteria.' This makes per-account features such as total post volume, active-day count, and observed timeline span systematically different between the classes by construction. The Related Work claim that the datasets include, for control accounts, 'other posts from their timelines' is inconsistent with this procedure. No quantification of the resulting distributional difference is provided, and Fig. 2 reports only aggregate collective coverage. Because a detection algorithm could separate the classes on this artifact alone, the paper should either release matched temporal subsets and per-class feature distributions, or explicitly restrict the benchmarking claim to controls for which the confounding is addressed.","section":"Discussion (limitations)"},{"comment":"The choice to define control accounts by the hashtags used by IO accounts makes the validity of the negative class depend on the quality of the hashtag proxy, as the paper itself acknowledges: 'the quality of the control sample depends on IO hashtags being a quality proxy for IO content, which is not always the case.' This is a genuine threat to the central contribution because campaigns with little or no hashtag use may lack valid control counterparts, and generic hashtags can attract non-IO coordinated or inauthentic activity into the control set. The paper gives an illustrative #crypto example but does not quantify the problem: per campaign, what fraction of IO posts contain no hashtagged content, and how specific are the hashtags used? Without such diagnostics, downstream users cannot tell which of the 26 campaigns support a reliable detection benchmark. Please add per-campaign hashtag-coverage statistics and, if possible, a control subset that excludes very generic hashtags.","section":"Discussion (limitations)"}],"minor_comments":[{"comment":"The phrase 'developed ntify messages' is truncated and should read 'developed to identify messages'.","section":"Related Work"},{"comment":"The sentence 'does not need require re-hydration' should be corrected to 'does not require re-hydration'.","section":"Related Work"},{"comment":"The sentence beginning 'Research who wish to mitigate' should read 'Researchers who wish to mitigate'.","section":"Discussion"},{"comment":"The claim 'Our datasets have 100% coverage of control data' should be qualified, since the collection procedure caps control timelines at 100 posts per day and excludes posts after the first matched date; the statement is stronger than what the data collection section describes.","section":"Related Work"}],"recommendation":"major_revision","confidential_remarks":"The manuscript does not explicitly name the social media platform that produced the original IO data, and the Ethical Checklist states the IO data are no longer public. Because provenance and redistribution rights are central to a dataset release, I would ask the editor to have the authors document the source platform and permission terms in the repository metadata, even if the platform is not named in the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is the real thing: 26 verified IO campaigns, full IO timelines, and over 13M control posts from 303k accounts, anonymized and ready without rehydration. That fills a concrete gap now that the Twitter API is effectively closed, and it is a step up from Guo and Vosoughi's 1% stream sample. I'd use it, and I'd bring it to a reading group.\n\nThe authors also deserve credit for how they disclose the warts. In the Discussion they explicitly say the control timelines are cropped at 100 posts and exclude post-inclusion activity, that hashtag-based selection can bias toward active accounts or pick up non-IO coordination, and that the quality of the control sample depends on hashtags being a good proxy. That's honest, and it matches what I see in the methods.\n\nThe soft spots are real, though. The stress-test note is right: for account-level detection, IO accounts contribute their entire timelines while control accounts contribute only single days capped at 100 posts. Any detector will trivially separate the classes on total post volume, active days, or account age. The authors suggest users 'selectively choose IO and control accounts with matching temporal activity,' but they don't ship such a subset or quantify how much the artifact alone separates the classes. That is load-bearing for the central benchmarking claim. Also, the Related Work says '100% coverage of control data,' which is contradicted by the 100-post cap; that's a minor phrasing issue, but in a dataset paper precision matters.\n\nI don't think the paper is fatally flawed. The control data are still valuable for descriptive studies and for researchers who apply the temporal matching themselves. But the datasets as released are not ready-to-benchmark without that extra handling, and the paper should be clearer that detection results on the raw data should not be interpreted as measuring IO-detection skill.\n\nWho is this for? People working on influence-operation detection or characterization who need a multi-country, multi-campaign resource. It deserves peer review — a serious referee can push for a temporal-matching appendix or a separate activity-matched release. I'd accept with revisions, not desk-reject.","headline":"A genuinely useful multi-campaign IO dataset with an honest limitations section, but the unquantified timeline asymmetry between IO and control accounts is a real confound for detection benchmarking.","tokens_in":11876,"tokens_out":1400,"would_cite":true,"duration_ms":16353,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper releases labeled datasets that pair platform-verified information-operation posts with over 13 million control posts from 303,000 accounts discussing the same topics at the same times, for 26 campaigns across 16 state actors.","keywords":["information operations","coordinated inauthentic behavior","social media datasets","control data","influence campaigns","detection benchmark","hashtag-based sampling","anonymized data"],"falsifier":"A concrete check would be to compute, per campaign, the fraction of IO posts that contain no hashtag and the fraction of IO hashtags that never appear in control posts; the paper's own numbers give a median hashtag coverage of only 31%, so campaigns whose distinctive hashtags are largely absent from the control data would yield biased benchmarks. One could also manually audit a random sample of control accounts for spam-like or coordinated patterns; any substantial share of non-organic control accounts would directly invalidate the negative labels.","tokens_in":10904,"feed_emoji":"🕵️","tokens_out":5197,"duration_ms":49280,"temperature":0.7,"pith_summary":"This paper introduces a public, anonymized dataset resource for studying state-backed information operations (IOs) on a major social media platform. For each of 26 verified campaigns attributed to 16 state actors, it pairs the platform-confirmed IO posts with a control set: over 13 million posts by about 303,000 accounts that discussed the same hashtags on the same dates and were not part of the operation. The paper argues that this control data was previously missing at this scale and that its absence blocked the development of detection algorithms that can generalize across campaigns and countries. Because the data is already hydrated and anonymized, researchers can use it directly without the platform APIs that are no longer accessible. If the datasets hold up, they give the field a common benchmark for characterizing and detecting coordinated inauthentic behavior.","feed_headline":"Verified IO posts plus organic controls for 26 campaigns","feed_subtitle":"Lets researchers benchmark influence-operation detection without re-hydrating tweets: control and IO data come ready-anonymized.","key_machinery":"The method that carries the paper is a hashtag-anchored control-selection pipeline. For each campaign, the authors extract all hashtags used by the IO accounts, query the platform API for accounts that used the same hashtags on the same dates, and then pull those control accounts' timelines for those dates, truncated at 100 posts per day. The working assumption is that hashtag co-use on the same day selects organic accounts engaging with the same topics at the same time, making 'is control' a usable negative label for supervised detection and a baseline for descriptive comparison. The pipeline also merges the IO and control records into a single schema with aligned field names and an explicit is_control column.","core_discovery":"The central delivery is a set of 26 labeled datasets, one per campaign, each with a binary flag indicating whether a post comes from a platform-verified IO account (false) or from a matched control account (true). The IO side contains the full timelines of the accounts the platform attributed to each campaign, not just the posts judged inauthentic; the control side contains daily timelines (up to 100 posts per day) of accounts that used the IO campaigns' hashtags on the same dates. The authors report 100% coverage of control data, unlike earlier control sets drawn from small samples, and they anonymize all identifiers with a consistent one-way hash so that mentions and reposts remain linkable across the IO/control boundary. Their claim is that this combination—verified malicious accounts, topic- and time-matched organic accounts, full timeline context, and cross-campaign breadth—supplies what detection research lacked: a reusable, re-hydration-free benchmark for distinguishing coordinated from organic behavior.","pith_inferences":["An implication the authors leave implicit is that the dataset's value is uneven across campaigns: because control selection depends on hashtags, campaigns whose IO accounts rarely hashtag (the paper reports hashtag coverage varying from 3% to 73%, with a median of 31%) will produce weaker controls, so benchmarks should report per-campaign results rather than pooled accuracy.","Since the anonymization scheme preserves account linkage across IO and control posts, a natural extension not pursued in the paper would be to release derived temporal interaction graphs, enabling direct evaluation of coordination-detection algorithms that operate on network structure.","Because the control labels are based on absence from a platform-verified IO list rather than on any independent audit, evaluation protocols that treat some control accounts as potentially noisy negatives—for example, by manually reviewing a sample or using robust ranking metrics—are more trustworthy than protocols that assume all control labels are clean."],"forward_implications":["Researchers can train supervised detectors to distinguish IO accounts from organic accounts across dozens of campaigns without needing to collect or re-hydrate posts through now-inaccessible platform APIs.","The multi-country, multi-actor structure allows cross-campaign generalization tests, such as training on some operations and evaluating on held-out campaigns from different states or languages.","Including full daily timelines for control accounts, rather than only topic-matched posts, enables behavior-sequence and account-history methods that earlier campaign-specific control sets did not support.","The consistent anonymization hash lets investigators reconstruct mention, repost, and reply networks across the IO/control boundary, supporting network-based coordination detection and descriptive studies of engagement tactics.","The datasets establish a shared benchmark where future detection methods can be compared against the same verified labels and control samples across 26 campaigns."],"supporting_citations":[{"why":"Introduced control datasets for 28 IO campaigns using the campaigns' top hashtags, the closest prior method this paper extends.","marker":"Guo and Vosoughi (2022)"},{"why":"Compiled control data for two IO campaigns by selecting top hashtags and collecting all tweets with those hashtags, the immediate methodological predecessor.","marker":"Cima et al. (2024)"},{"why":"Collected tweets based on hashtags and keywords related to the 2016 U.S. election to form a control set for the Russian IRA campaign.","marker":"Badawy et al. (2019)"},{"why":"Documents the shutdown of the Twitter API that motivates providing hydrated, re-hydration-free data.","marker":"Murtfeldt et al. (2024)"},{"why":"A detection method based on sequences of account actions that requires control data of the kind these datasets provide.","marker":"Nwala, Flammini, and Menczer (2023)"},{"why":"Detects accounts from previously unseen campaigns using cross-campaign data, which needs labeled multi-campaign datasets to train and evaluate.","marker":"Saeed et al. (2024)"},{"why":"Curated control datasets for four campaigns by querying keywords, providing a comparison baseline for control-data construction.","marker":"Smith, Ehrett, and Warren (2024)"}],"fun_headline_variants":["26 labeled IO campaigns with matched organic controls","IO posts + organic controls for 26 campaigns","Benchmark IO detection with verified and control data","13M control posts to test IO detection algorithms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The control labels are only as good as the assumption that the hashtags used by IO accounts capture the content of each campaign, and that accounts sharing those hashtags on the same dates are genuinely organic—an assumption the paper itself notes fails when IO tactics avoid hashtags or when generic hashtags pull in spammers and other coordinated actors.","fun_headline_variants_meta":{"raw":{"variants":["26 labeled IO campaigns with matched organic controls","IO posts + organic controls for 26 campaigns","Benchmark IO detection with verified and control data","13M control posts to test IO detection algorithms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000479,"raw_usage":{"total_tokens":2335,"prompt_tokens":875,"completion_tokens":1460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":1402}},"tokens_in":491,"tokens_out":1460,"duration_ms":12298,"temperature":1.0,"reasoning_tokens":1402,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:30:35.444904+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check would be to compute, per campaign, the fraction of IO posts that contain no hashtag and the fraction of IO hashtags that never appear in control posts; the paper's own numbers give a median hashtag coverage of only 31%, so campaigns whose distinctive hashtags are largely absent from the control data would yield biased benchmarks. One could also manually audit a random sample of control accounts for spam-like or coordinated patterns; any substantial share of non-organic control accounts would directly invalidate the negative labels.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced control datasets for 28 IO campaigns using the campaigns' top hashtags, the closest prior method this paper extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Compiled control data for two IO campaigns by selecting top hashtags and collecting all tweets with those hashtags, the immediate methodological predecessor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Collected tweets based on hashtags and keywords related to the 2016 U.S. election to form a control set for the Russian IRA campaign."},{"cited_title":"C.; Flammini, A.; and Menczer, F","cited_arxiv_id":null,"evidence_quote":"A detection method based on sequences of account actions that requires control data of the kind these datasets provide."},{"cited_title":"Unsupervised detection of coordinated information operations in the wild","cited_arxiv_id":"2401.06205","evidence_quote":"Curated control datasets for four campaigns by querying keywords, providing a comparison baseline for control-data construction."}],"review_version":1}