{"id":"008e1e18-72b9-40ec-a1d3-089078d16c1e","arxiv_id":"2505.24724","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Ethereum's transaction input data field is used as a decentralized messaging medium, with English messages dominated by security warnings and Chinese messages by emotional expression.","lead":"Researchers analyzed 867,140 Ethereum transactions whose input data fields contain readable text messages, finding that people use the blockchain as a messaging channel. The messages split sharply by language: English text is mostly security warnings and negative emotions, while Chinese text is mostly personal and emotional expression.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central cultural-divergence percentages rest on unvalidated GPT-4o labels; a human-annotation validation set is needed to confirm the English/Chinese topic and emotion split.","rationale":"The paper opens a novel empirical object and its qualitative examples are plausible; we are not accusing the authors of any misconduct. The load-bearing premise for the headline contribution is that the LLM labels used for topic and sentiment are sufficiently accurate to support the claimed cross-language divergence. That premise is unverified: no gold-standard evaluation is reported, and the prompt definitions plausibly nudge the model toward the reported pattern. I considered the reply-rate analysis (§8.1) as an alternative, but that is a secondary finding and its definition of 'reply' is not even stated; the cultural-divergence percentages are the abstract's central quantitative claim. I also considered the EOA-only scope, but the paper explicitly frames IDMs as the EOA-to-EOA, UTF-8-decodable subset, so that is a defined scope rather than a hidden assumption. The proposed human-annotation check would settle whether the concern lands; if the human-labeled distributions match, the claim is supported, and if not, the percentages need revision. This aligns with the reader's conditional verdict, so the verdict is unchanged.","tokens_in":24484,"tokens_out":4969,"duration_ms":52825,"concrete_test":"Randomly sample 500 English and 500 Chinese IDMs from the natural-language subset. Have two independent bilingual annotators, blind to GPT-4o outputs, label each message with the same topic and emotion taxonomy. Compute Cohen's kappa between annotators and between each annotator and GPT-4o; compare the human-labeled English vs Chinese topic/emotion distributions to the reported ones. If kappa < 0.6 for either language, or if the human-labeled Chinese Social & Emotional Expression share is below 30% (vs 44%) or the English Security & Incident share differs by more than 10 percentage points, the cultural-divergence claim is not robust to labeling error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that English IDMs center on Security & Incident (24%) with predominantly negative emotions while Chinese IDMs center on Social & Emotional Expression (44%) with positive tone—is computed from GPT-4o assignments (Table 3, Figures 5–6, Table 4). No validation set, inter-annotator agreement, or error rates are reported (§5.1.1, §5.2, §12). The Appendix A prompt defines categories in ways that can steer labels: Negative-Fear explicitly includes 'caution about risks, scams, or vulnerabilities' and Security & Incident includes 'Warnings on malicious activities'; if the model uses these cues, English security messages may mechanically map to Fear, and Chinese relational phrases (e.g., '520') may be tagged as Love/Joy without independent evidence. Although the model returns confidence scores, no threshold is applied, so low-confidence labels contribute equally. The paper's own limitations acknowledge LLM prompt sensitivity but offer no quantitative check. Without a labeled gold set, the reported English/Chinese divergence could be an artifact of the taxonomy/prompt rather than a property of the IDM corpus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies human-readable messages embedded in the input data field of Ethereum transactions, which the authors call Input Data Messages (IDMs). Using a crawl from the genesis block to February 2024, the authors identify 5,238,336 UTF-8-decodable IDM transactions and, after filtering, 867,140 informative IDMs from EOA-to-EOA transactions. They use GPT-4o for language detection, topic classification (12 main topics, 48 subtopics), and sentiment analysis (16 emotions), then report descriptive statistics on topic and emotion distributions for English and Chinese, value/length patterns, network structure, security-related victim-to-attacker messages, and toxic content. The central claim is that Ethereum's input data field functions as a decentralized communication medium with culturally patterned usage: English messages focus on security warnings and negative emotions, while Chinese messages emphasize social and emotional expression with positive tone.","tokens_in":24679,"tokens_out":7266,"duration_ms":76027,"significance":"If the semantic findings are accurate, the paper opens a new empirical area: a large, previously unmeasured social messaging layer on Ethereum, and it documents cultural divergence that is of interest to both blockchain researchers and social computing. The paper's contributions include a large crawled dataset, a transparent decoding and filtering pipeline, and a useful taxonomy for on-chain messages. The authors also acknowledge some limitations and take steps to avoid re-publishing toxic content. However, the central percentages rest entirely on unvalidated LLM labels, and the security-effectiveness claims lack a baseline, so the quantitative conclusions are not yet reliable. With the addition of a human-validated gold set and a proper reply-rate baseline, the paper could become a solid reference for this new area.","major_comments":[{"comment":"The paper's headline findings—English IDMs centered on Security & Incident (24%) and Chinese on Social & Emotional Expression (44%), with contrasting negative/positive emotional tones—are computed exclusively from GPT-4o labels assigned in §5.1 and §5.2. No validation set, inter-annotator agreement, per-class precision/recall, or error rates are reported anywhere, and the limitations section (§12) only asserts that prompts were carefully designed and that human review was used. Because the Appendix A prompt defines Negative-Fear as 'caution about risks, scams, or vulnerabilities' and the Security & Incident topic includes 'Warnings on malicious activities,' the observed alignment between English security messages and Fear could be an artifact of the taxonomy's definitions. In addition, the sentiment classification collects a confidence score but applies no threshold, so low-confidence labels weight equally with high-confidence ones. A human-annotated gold set (at least a few hundred messages per language), with reported agreement or error metrics, is required before the cultural-divergence percentages can be considered measurements of the corpus rather than properties of the model.","section":"§5.1.1, §5.2, §12 (Tables 3–4, Figs. 5–6)"},{"comment":"The claim that negotiation and reward offers are 'linked to higher reply rates' (15.9% and 19.6%, vs. 7.3% for pleading and 5.3% for threatening) is not supported by the reported analysis. There is no baseline reply rate for the pool of security-related IDMs as a whole, nor for a suitable control group (e.g., other IDM types), and there is no definition of what counts as a 'reply' (e.g., a subsequent transaction from the attacker address to the victim within a time window). Without these, one cannot tell whether the observed differences exceed the general propensity of attackers to respond to any incoming transaction, or whether they are statistically significant given the small sample sizes (n=245 and n=485 for Reward and Negotiate).","section":"§8.1, Table 5"},{"comment":"The paper frames itself as a study of 'Ethereum IDMs' from the genesis block onward, but the actual analyzed dataset is a filtered subset: §3 keeps only EOA-to-EOA transactions whose input data decodes as UTF-8, excluding all contract-directed ABI-encoded calls and any encoded/encrypted payloads. The abstract's '87%+ historical transactions' refers to the crawled block range, not to the 867,140 informative IDMs that remain after filtering; this is misleading about the breadth of the analysis. Because contract-mediated messages and non-UTF-8 encodings may carry different linguistic content, the paper's conclusions about 'what Ethereum users talk about' should be explicitly scoped to decodable EOA-to-EOA IDMs.","section":"Abstract, §2, §3"}],"minor_comments":[{"comment":"The number of senders is inconsistent: §1 reports 59,795 senders for all informative IDMs, while §7.1 states that 189,111 addresses send the 422,387 natural-language IDMs. Please clarify the filter for each number or correct the error.","section":"§1 vs. §7.1"},{"comment":"The counts of dyadic communities are inconsistent: the text first reports 15,625 communities of size two, then refers to 'the 15,205 smallest communities of exactly two addresses.' Please reconcile these numbers.","section":"§7.2"},{"comment":"The 'Percentage' column in Table 3 appears to be aggregate across languages, while Figures 5 and 6 show language-specific shares; add a sentence to the caption explaining the base used for each percentage, since the current layout can make 18.7% and 24% for Security & Incident look contradictory.","section":"Table 3 caption"},{"comment":"The cumulative percentage curve would be easier to interpret if the caption indicated that it corresponds to the right-hand y-axis, and if the '80% of unique IDMs are shorter than 100 bytes' claim were marked on the plot.","section":"Figure 4"},{"comment":"The paper should state the number of samples subjected to human review and how disagreements with GPT-4o were resolved; the current description ('human review and iterative prompting adjustments') is too vague to assess classification reliability.","section":"§10 and §12"},{"comment":"The strategies Plead, Threaten, Reward, and Negotiate are identified by LLM; the same validation concern applies here as in §5, and at minimum a few examples per strategy should be shown so readers can judge the labeling.","section":"§8.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for a security/cryptocurrency-empirical venue and will likely be cited, but the headline numbers need stronger support. I would not reject, yet I would insist that the validation set and reply-rate baseline be added prior to acceptance. The authors should also be asked to reconcile the sender-count and community-count inconsistencies, as they currently undermine trust in the network analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, know this: the paper is a real first. It treats Ethereum's input data field as a social communication channel and measures it at scale. Nobody has done that. The descriptive material—867k informative IDMs, the English/Chinese topic split, the 520/1314 love payments, the warning-message culture, the toxic-content sample—is genuinely new and mostly credible as observation. The network and community analysis is standard but competently done, and the authors are transparent about their EOA-to-EOA and UTF-8-decodable filters.\n\nThe centerpiece is the cultural divergence: English IDMs are security-heavy and fearful, Chinese IDMs are social/emotional and positive. That result is plausible, but it is computed entirely from GPT-4o labels, and the paper reports no validation set, no inter-annotator agreement, and no error rates. The Appendix A prompt makes this more than a generic concern: Negative-Fear is defined as \"caution about risks, scams, or vulnerabilities,\" and the Security & Incident topic includes \"Warnings on malicious activities.\" The model is cued to produce exactly the mapping the paper then reports. Low-confidence labels are accepted equally. The limitations section mentions human review and prompt adjustment, but gives no procedure or quantity. Before I would trust the 24% vs 44% split, I would need a labeled gold set of a few hundred messages per language with human labels and GPT-4o agreement rates.\n\nOther soft spots are smaller but real. The reply-rate table compares four recovery strategies but has no baseline, no significance testing, and small counts; that claim should be read as suggestive, not established. The EOA-to-EOA filter is transparent, but it means the measured \"communication layer\" excludes contract-directed messages and encoded payloads, so the title's scope is narrower than the abstract implies. And for an empirical paper, the dataset and code being available only \"upon request\" is a genuine gap.\n\nThe descriptive direction holds: people do use input data to talk, and the English/Chinese difference is unlikely to be pure noise. But the exact percentages are not yet load-bearing. This paper is for blockchain socio-technical researchers, security people studying scam warnings, and anyone working on on-chain content governance. It deserves a serious referee, with the central revision being a validation study. I would accept it for review, and I would cite it for the phenomenon and the descriptive patterns—but not for the specific cultural-divergence percentages until the validation appears.","headline":"A genuine first measurement of Ethereum's input-data chatter, with a cultural-divergence centerpiece that depends on unvalidated LLM labels—worth refereeing, but the percentages need a human-labeled validation set before they are load-bearing.","tokens_in":25193,"tokens_out":2035,"would_cite":true,"duration_ms":27216,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ethereum's input data field is a real messaging channel: 867,140 informative messages reveal distinct English and Chinese on-chain cultures.","keywords":["Ethereum","input data messages","on-chain communication","semantic analysis","sentiment analysis","network analysis","content moderation","large language models"],"falsifier":"Take a stratified random sample of, say, 500 English and 500 Chinese IDMs, have independent human coders label language, topic, and emotion using the paper's taxonomy without seeing the LLM output, and compute agreement; if the LLM's per-class precision or recall is materially below the level needed to sustain the 24% versus 44% topic gap, or if the errors correlate with language, the cultural-divergence result is an artifact of labeling bias rather than a property of the messages.","tokens_in":24295,"feed_emoji":"💬","tokens_out":8041,"duration_ms":86970,"temperature":0.7,"pith_summary":"The paper establishes that Ethereum's input data field, designed for smart-contract calls, is actively repurposed as a peer-to-peer messaging channel. Analyzing transactions from the genesis block to February 2024, it isolates 867,140 informative messages embedded in ordinary user-to-user transfers. Its central substantive claim is cultural: English-language messages concentrate on security warnings and scam alerts with predominantly negative emotion, while Chinese-language messages center on social and emotional expression with a positive tone. The paper also claims that IDM participants form small, loosely connected communities, that victims use the channel to negotiate with attackers, and that unmoderated toxic content on-chain creates a governance gap. If the analysis is right, Ethereum carries a substantial social layer that standard financial readings of blockchain data miss.","feed_headline":"Ethereum transactions talk: 867,140 on-chain messages found","feed_subtitle":"English IDMs warn about scams and hacks; Chinese IDMs share love and daily life, a first large-scale study shows.","key_machinery":"The central object is the Input Data Message (IDM): human-readable text embedded in the input data field of transactions between ordinary user accounts (externally owned accounts, EOAs), decoded from UTF-8. The argument is carried by a filtering pipeline that reduces over five million decodable transactions to 867,140 informative IDMs, followed by a large-language-model-based classifier using a hand-built taxonomy of 12 topics, 48 subtopics, and 16 emotion categories, plus a community-detection pass over the sender-receiver graph. The taxonomy is the instrument that turns raw hex into the English-versus-Chinese cultural comparison, and the graph analysis turns addresses into statements about broadcast-heavy versus personal communication patterns.","core_discovery":"The discovery is that the input data field has become a used communication medium, not just an ABI payload, and that this on-chain talk is culturally patterned. English and Chinese account for 99.8% of natural-language IDMs, with English at 95.4% and Chinese at 4.4%. English IDMs cluster in Security & Incidents (24%), led by warnings about phishing and rug pulls (17%), and Fear is the dominant emotion, with 196,338 total occurrences. Chinese IDMs cluster in Social & Emotional Expression (44%), led by love and confession (18%), daily life records (8%), and birthday wishes (7%), with Joy and Love dominant, and 24% of Chinese IDMs are on-chain copyright certificates. Longer English IDMs tend to accompany high-value ETH transfers for protocol-level purposes, while longer Chinese IDMs often carry symbolic amounts such as 5.201314 ETH for emotional intent. The IDM network is mostly one-directional: the average clustering coefficient is 0.01, 59.99% of communities consist of exactly two addresses, and the largest community issues 34.9% of all IDMs. In security incidents, recovery messages that offer rewards or propose negotiation draw reply rates of 15.9% and 19.6%, versus 7.3% for pleading and 5.3% for threatening.","pith_inferences":["Editorial inference: if the LLM labels hold, the quarterly topic time series functions as an on-chain public-sentiment index, since spikes visibly track events like the 2018 #MeToo wave and COVID-19 lockdown discussions.","Editorial inference: the reply-rate comparison points to a testable practical playbook, namely that victims who open with a reward or compromise proposal rather than a threat are more likely to get an attacker response; a prospective analysis of post-2024 incidents could test this.","Editorial inference: the EOA-to-EOA, UTF-8-decodable filter excludes contract-directed inputs and encoded payloads, so the 99.8% English-plus-Chinese split describes one visible slice of the channel; a broader definition including contract calls would likely shift the topic mix and could dilute the cultural divergence."],"forward_implications":["Ethereum's ledger is a permanent, public archive of user speech, so future social-science work can treat on-chain messages as historical records that cannot be edited or removed.","English-speaking IDM users predominantly use the channel to warn about scams and hacks, while Chinese-speaking users use it for love confessions, daily-life records, and symbolic transfers, implying that communication norms, not just costs, shape blockchain usage.","Longer messages correlate with different purposes by language: high-ETH protocol transfers in English versus emotionally meaningful micro-amounts in Chinese, which matters for interpreting transaction-value statistics.","The IDM network is broadcast-heavy rather than conversational, with 95.5% of addresses engaged in one-directional messaging and only two-address communities dominating the periphery, so on-chain talk resembles public bulletin boards more than chat rooms.","Victim-to-attacker fund-recovery messages with rewards or negotiation draw materially higher reply rates than pleading or threatening, suggesting a measurable, strategy-dependent negotiation channel on-chain."],"supporting_citations":[{"why":"Defines the input data field and the concept of an Input Data Message that the paper analyzes.","marker":"[4]"},{"why":"Defines Ethereum's transaction structure, including input data as the field repurposed for messaging.","marker":"[2]"},{"why":"Provides the original peer-to-peer electronic cash framing that the paper extends to peer-to-peer communication.","marker":"[1]"},{"why":"Documents limitations of automatic language identification, motivating the LLM-based refinement of language labels.","marker":"[6]"},{"why":"Supplies the 2019 gas-cost reduction used in the analysis of why IDM storage costs fell but adoption rose only later.","marker":"[11]"},{"why":"Supplies the community-detection algorithm used to find the 26,048 IDM communities and their sizes.","marker":"[12]"},{"why":"Provides the on-chain attack taxonomy that frames the security-relevance analysis of warnings and fund-recovery IDMs.","marker":"[13]"},{"why":"Supplies Tornado Cash depositor addresses used to test overlap between security-related IDM users and privacy-seeking behavior.","marker":"[14]"},{"why":"Documents content moderation shortcomings in web3 platforms, used to frame the moderation and regulation implications.","marker":"[16]"}],"fun_headline_variants":["Ethereum transactions hide chats: 867K messages found","English warns, Chinese emotes: 867K on-chain messages","Ethereum's input data is a chat channel: 867K messages","Blockchain talk: 867K messages, security vs emotion","Ethereum IDMs: 867K on-chain messages, culturally split"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every language, topic, and emotion percentage in the paper rests on labels produced by a single large language model, reviewed only qualitatively, with no validation set or measured error rate; if those labels are systematically biased by language or prompt wording, the central cultural-divergence claim collapses.","fun_headline_variants_meta":{"raw":{"variants":["Ethereum transactions hide chats: 867K messages found","English warns, Chinese emotes: 867K on-chain messages","Ethereum's input data is a chat channel: 867K messages","Blockchain talk: 867K messages, security vs emotion","Ethereum IDMs: 867K on-chain messages, culturally split"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2489,"prompt_tokens":1168,"completion_tokens":1321,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":784,"completion_tokens_details":{"reasoning_tokens":1228}},"tokens_in":784,"tokens_out":1321,"duration_ms":13010,"temperature":1.0,"reasoning_tokens":1228,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:14:22.178922+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of, say, 500 English and 500 Chinese IDMs, have independent human coders label language, topic, and emotion using the paper's taxonomy without seeing the LLM output, and compute agreement; if the LLM's per-class precision or recall is materially below the level needed to sustain the 24% versus 44% topic gap, or if the errors correlate with language, the cultural-divergence result is an artifact of labeling bias rather than a property of the messages.","supporting_citations":[{"cited_title":"Etherscan information center: Understanding transaction input data","cited_arxiv_id":null,"evidence_quote":"Defines the input data field and the concept of an Input Data Message that the paper analyzes."},{"cited_title":"Ethereum: A secure decentralised generalised transaction ledger","cited_arxiv_id":null,"evidence_quote":"Defines Ethereum's transaction structure, including input data as the field repurposed for messaging."},{"cited_title":"Bitcoin: A peer-to-peer electronic cash system","cited_arxiv_id":null,"evidence_quote":"Provides the original peer-to-peer electronic cash framing that the paper extends to peer-to-peer communication."},{"cited_title":"Automatic language identification in texts: A survey.Journal of Artificial Intelligence Research, 65:675–782, 2019","cited_arxiv_id":null,"evidence_quote":"Documents limitations of automatic language identification, motivating the LLM-based refinement of language labels."},{"cited_title":"EIP-2028: Transaction data gas cost reduction (settled)","cited_arxiv_id":null,"evidence_quote":"Supplies the 2019 gas-cost reduction used in the analysis of why IDM storage costs fell but adoption rose only later."},{"cited_title":"Fast unfolding of communities in large networks","cited_arxiv_id":null,"evidence_quote":"Supplies the community-detection algorithm used to find the 26,048 IDM communities and their sizes."},{"cited_title":"Sok: Decentralized finance (DeFi) attacks","cited_arxiv_id":null,"evidence_quote":"Provides the on-chain attack taxonomy that frames the security-relevance analysis of warnings and fund-recovery IDMs."},{"cited_title":"On how zero-knowledge proof blockchain mixers improve, and worsen user privacy","cited_arxiv_id":null,"evidence_quote":"Supplies Tornado Cash depositor addresses used to test overlap between security-related IDM users and privacy-seeking behavior."},{"cited_title":"Under- standing and improving content moderation in web3 platforms","cited_arxiv_id":null,"evidence_quote":"Documents content moderation shortcomings in web3 platforms, used to frame the moderation and regulation implications."}],"review_version":1}