Pith. sign in

REVIEW 4 major objections 5 minor 86 references

A 2.98-billion-token corpus spanning science, patents, and social media shows that DLT concepts first appear in research before markets.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-02 20:46 UTC pith:AGLSTIY6

load-bearing objection The corpus is a real resource worth having; the diffusion and market-lead analyses overreach, but that's fixable. the 4 major comments →

arxiv 2602.22045 v2 pith:AGLSTIY6 submitted 2026-02-25 cs.CL

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

classification cs.CL
keywords distributed ledger technologyblockchaintext corpusnamed entity recognitionsentiment analysisinnovation diffusionpatent analysisdomain-adapted language model
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper introduces DLT-Corpus, a publicly released collection of 2.98 billion tokens from 22.12 million documents spanning scientific publications, US patents, and social media posts. It argues that this is the largest domain-specific text resource for distributed ledger technology and demonstrates its value by showing that technologies like stablecoins, decentralized exchanges, and automated market makers first appear in scientific literature, then in patents, then in social media. It also reports that scientific publication volume leads cryptocurrency market expansion by about two years, while social media sentiment stays overwhelmingly bullish even during market downturns. Additionally, a language model trained on this corpus improves named-entity recognition for DLT terms by 23% over a general baseline while preserving general sentiment performance.

Core claim

The central claim is that DLT-Corpus, the largest domain-specific text collection for distributed ledger technology, enables both specialized NLP and cross-domain studies of innovation. Using it, the authors find that three economically significant technologies (stablecoins, DEXs, AMMs) consistently appear in scientific literature before they appear in patents or social media, matching a traditional technology-transfer pattern. They also find an asymmetric lagged correlation: scientific publications lead market capitalization by two years (rho=0.95, p<0.001) and the correlation decays when the market leads, while patents show a symmetric correlation and social media follows the market. The c

What carries the argument

The central object is the corpus itself, which aggregates three timestamped document streams: scientific literature (37,440 open-access publications), US patents (49,023 records), and social media posts (22 million). The analysis relies on the corpus's temporal metadata to compute lagged correlations and first-mention dates, and on its keyword density (8.7 times higher than general web corpora) to argue it provides concentrated learning signal for domain-adapted models. The filtering pipeline, including a BERT-based domain relevance filter trained on an existing DLT NER dataset and manual pruning, defines the corpus's membership and thus drives the first-mention ordering.

Load-bearing premise

The technology-transfer ordering rests on first-mention dates derived from a corpus whose scientific subset was filtered by a BERT-based relevance model and whose patents were selected by a simple keyword search; different query or threshold choices could change the order.

What would settle it

For any of the three analyzed technologies (stablecoins, DEXs, AMMs), find a patent or social media post dated earlier than the corpus's first scientific mention of that term; such a document would break the claimed science-first ordering.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Tracking scientific literature can serve as an early-warning system for emerging DLT technologies, since concepts first appear there.
  • Scientific output appears to lead cryptocurrency market growth by about two years, suggesting research is a leading indicator for DLT market expansion.
  • Domain-adapted language models trained on DLT-Corpus outperform general models on DLT-specific NER, showing concentrated terminology matters.
  • The corpus enables integrated analysis across scientific, patent, and social discourse that previously required assembling separate, often inaccessible datasets.
  • Because social media coverage stops at 2023, the corpus is a fixed snapshot and cannot track post-2023 community discourse.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The first-mention ordering (science before patents before social media) is sensitive to how each subset was collected and filtered; a differently filtered corpus could yield a different diffusion ordering, so the technology-transfer conclusion should be tested on alternative query and filtering choices.
  • The two-year lead of scientific publications over market cap is based on annual aggregates; finer-grained data could reveal shorter lags or lead-lag dynamics within quarters, and the correlation alone does not establish causation.
  • The persistently bullish social sentiment may reflect the cryptocurrency-focused composition of the social media sources and the sentiment labeling method rather than an objective market mood; a sample of general social media would be a useful check.
  • The corpus excludes news articles for copyright reasons; a news stream, if added, might show that financial journalism acts as an intermediary between research and market sentiment.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper introduces DLT-Corpus, a large multi-source text collection for the Distributed Ledger Technology domain, containing 2.98 billion tokens from 37,440 scientific publications, 49,023 USPTO patents, and 22.03 million Twitter posts. The authors describe the construction and quality assessment of the corpus, train a domain-adapted language model (LedgerBERT), release a crowdsourced sentiment-analysis dataset, and present two utility analyses: (i) correlations between document volumes, market capitalization, and sentiment; and (ii) technology diffusion across scientific literature, patents, and social media. The headline claims are that technologies originate in scientific literature before reaching patents and social media, and that scientific publications lead market expansion by two years (ρ=0.95, p<0.001). The resource is publicly released with code, models, and datasheet-style documentation.

Significance. If the claims hold, DLT-Corpus is a substantial and valuable contribution: it is described as the largest domain-specific DLT text collection, it is publicly released, and it ships reproducible artifacts (code, models, datasets) as well as a FAIR-aligned datasheet. The quality assessment against general-purpose corpora is a useful benchmark. The temporal and cross-source analyses are of interest to innovation-diffusion and bibliometric communities. However, the two central analytical claims — technology-transfer ordering and science-leading-market causality — are currently supported only by correlations and normalized mention proportions, and the evaluation of LedgerBERT is partly in-domain by construction. The resource itself is likely useful even if those analytical claims are weakened; the paper's contribution should be reframed accordingly.

major comments (4)
  1. [§6.1, Figs. 3–4] The central claim that technologies 'originate in scientific literature before reaching patents and social media' is not established by the reported evidence. Figures 3 and 4 show normalized yearly proportions, not first-mention dates per technology or per document, and no statistical test for ordering is provided. More importantly, the three subsets have different temporal coverage (scientific: 1978–2025; patents: 1990–2025; social media: 2013–mid 2023; §4.1.3) and different collection filters. The patents subset uses only the search terms 'Distributed Ledger Technology' and 'blockchain' (§4.1.2), which can miss early patents using terms such as 'cryptocurrency', 'digital asset', or 'smart contract', biasing patent first-mentions late. The social media subsets are aggregated from Kaggle and academic datasets whose earliest coverage and sampling gaps are not documented; if those datasets
  2. [Abstract; §6, Tables 4–5] The abstract states that 'scientific publications lead market expansion by two years (ρ=0.95, p<0.001)', but Table 5 shows that social media also exhibits ρ=0.95 at lag −3, and patents exhibit ρ=0.95 at lag −2 and ρ=0.93 at lag −3. The claim that scientific literature uniquely or particularly leads is not supported by any test for differences among lags or among document types. Moreover, the correlations are computed on annual data from 2013–2024 (n=12 for scientific literature and patents, n=11 for social media) with no correction for autocorrelation or multiple testing; Spearman's ρ on such short, highly autocorrelated series can be large and significant even when no true lead–lag relationship exists. The full text also reports concurrent ρ=0.76 for scientific literature (Table 4), so the 'two-year lead' claim rests on a single lagged correlation without a formal comparison or robustne
  3. [§4.1.1 and §5.1] The primary evaluation of LedgerBERT is partly circular. The scientific-literature subset of DLT-Corpus was filtered by fine-tuning BERT-base-cased on the NER dataset from [36] and predicting domain entities (§4.1.1). LedgerBERT is then evaluated on exactly that same NER dataset from [36] as the 'primary evaluation' (§5.1). This means the model is trained on a corpus that was selected using labels from the very evaluation set, and the reported +23% over BERT-base and +3.5% over SciBERT may be inflated by construction. The authors note that the dataset 'derives from scientific literature, matching our corpus composition', but they do not discuss the selection overlap. An independent held-out evaluation (e.g., NER on patents or social media, or a DLT NER dataset not used in corpus filtering) is needed to support the claim that LedgerBERT improves DLT-specific NER because of domain-adaptive
  4. [§9 Limitations] The limitations section is incomplete regarding the two headline analyses. It states that 'marginally relevant DLT papers may remain in the dataset' but does not acknowledge that the science-first diffusion order and the market-lead correlations could be artifacts of (i) the patents subset's narrow keyword selection, (ii) undocumented temporal coverage or gaps in the social-media aggregations, or (iii) the scientific-filtering model's vocabulary biases. Because these threats bear directly on the paper's central claims, they should be addressed explicitly, preferably with supplementary analyses that restrict all three subsets to a common overlapping period and/or use alternative keyword vocabularies.
minor comments (5)
  1. [§6.1 text] The text refers to 'Fig. 2a' for stablecoins, 'Fig. 2b' for DEXs, and 'Fig. 2c' for AMMs, but the figures are numbered Fig. 3(a)–(c). Please correct the cross-references.
  2. [§6.2] The sentence 'Comparing Fig. 1d with Fig. 5' appears to refer to Fig. 1(d), the temporal evolution panel, but the connection between sentiment and document growth is asserted rather than quantified. Consider reporting a correlation between sentiment fractions and document volumes, or clearly labeling this as a visual comparison.
  3. [§4.2] The description of the sentiment-label construction would benefit from more detail: the 'median minimum votes' filter and the 25th/75th percentile boundaries are mentioned, but the exact thresholds and the number of examples excluded at each step are not reported. This is a reproducibility concern for the sentiment dataset.
  4. [Appendix A, Table 8] The social-media subset is described as 'Twitter/X posts', but several upstream sources are Kaggle datasets (e.g., 'bitcoin-tweets', 'crypto-tweets'). It would be useful to document which sources cover which time ranges, since this directly affects the diffusion analysis.
  5. [Throughout] Minor typographical issues: 'cryptocurrencies price prediction' in the abstract; 'T&Cs at collection time' should be 'terms and conditions'; the footnote on line 30 of §5.2 is incomplete. A careful proofread is recommended.

Circularity Check

1 steps flagged

LedgerBERT's flagship NER gain is measured on the same [36] annotations used to filter its own pretraining corpus; the rest of the claims are largely self-contained.

specific steps
  1. other [§4.1.1 (Domain filtering) and §5.1 (Primary evaluation: in-domain NER), Table 2]
    "To ensure relevance, we fine-tuned BERT-base-cased on the NER dataset from [36] and predicted domain-specific entities in each document. // NER serves as the primary evaluation of corpus quality because performance directly reflects how well the model learned domain-specific terminology. We use the DLT-focused NER dataset from [36]... This dataset derives from scientific literature, matching our corpus composition."

    The same authors' [36] NER annotations are load-bearing twice: they train the BERT filter that selects which scientific documents enter DLT-Corpus (§4.1.1), and they are the benchmark on which LedgerBERT is evaluated (§5.1) after continued pretraining on that filtered corpus. The reported +23.05% NER gain over BERT-base is therefore measured on a label scheme that already shaped the training corpus: documents were retained only if a model trained on [36] entities scored them highly. The evaluation is not a mathematical identity (the model must still learn the mapping), but as "the primary evaluation of corpus quality" it partly measures the same signal used in construction, so the utility claim is partly self-referential rather than independently validated.

full rationale

The corpus itself, its size, its public release, and the external keyword-density comparison (Table 1) are not defined in terms of the paper's conclusions, so the core dataset contribution is self-contained. The innovation-diffusion and lagged-correlation analyses are empirical exercises on the released artifacts; they may be confounded by unequal temporal coverage and keyword-based patent selection (a correctness risk the authors' §9 does not address), but no equation reduces the observed ordering to the paper's inputs. The one clear circularity is the [36] filter/evaluation loop: LedgerBERT is pretrained on a scientific subset selected by a BERT model fine-tuned on [36]'s NER labels, and then the paper's headline utility result is F1 on that same [36] NER dataset. Because [36] is the authors' own prior work and is invoked both for corpus construction and as the primary benchmark, the evaluation is in-domain by construction. Yet the model's gain is small and not logically forced, and the out-of-domain sentiment benchmark and external corpus comparisons provide independent content, so the appropriate score is 4 rather than a higher one.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 0 invented entities

The corpus itself and its derived statistics are the contribution. No new physical entities or theoretical constructs are proposed. The main unverified ingredients are representativeness assumptions and filtering thresholds that directly affect the innovation-diffusion timing claims.

free parameters (3)
  • Domain-filtering BERT thresholds = max prediction score > 0.995, or median score at/above subset-wide median
    Chosen by the authors to balance precision and coverage; directly determines which papers enter the scientific corpus and thus the first-mention findings.
  • Sentiment label percentile boundaries = 25th and 75th percentiles of normalized vote percentages
    Chosen boundaries for bullish/bearish/neutral; different cutoffs change sentiment time series.
  • Manual removal of 570 papers = 570 papers removed
    Hand-selected exclusion; if the wrong papers were removed, the diffusion timing could shift.
axioms (4)
  • domain assumption Semantic Scholar open-access papers with the DLT queries form a representative sample of DLT scientific literature.
    Query and database selection defines the scientific subset; not justified against e.g. Scopus/DBLP.
  • domain assumption Keyword-search patents with 'Distributed Ledger Technology'/'blockchain' capture the relevant patent landscape.
    Simple OR-style query; no recall/precision evaluation. Missing e.g. patents describing the tech without those literal keywords.
  • domain assumption Aggregated Kaggle/industry tweet collections are a representative snapshot of DLT community discourse.
    Selection of upstream tweet sets and pre-2023 cutoff shapes social-media trends; reweighting could change the correlation and sentiment results.
  • domain assumption First-mention dates in the corpus reflect when technologies actually first appeared in each community.
    The diffusion conclusion assumes that the corpus's earliest mention is close to the true origin; OCR errors, missing papers, and tweet collection gaps break this assumption.

pith-pipeline@v1.3.0-alltime-deepseek · 20363 in / 6301 out tokens · 49979 ms · 2026-08-02T20:46:59.986391+00:00 · methodology

0 comments
read the original abstract

We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents spanning scientific literature (37,440 publications), United States Patent and Trademark Office (USPTO) patents (49,023 filings), and social media (22 million posts). Existing Natural Language Processing (NLP) resources for DLT focus narrowly on cryptocurrency price prediction and smart contracts, leaving domain-specific language underexplored despite the sector's ~$3 trillion market capitalization and rapid technological evolution. We demonstrate DLT-Corpus' utility by analyzing patterns of technology emergence and market-innovation correlations. Findings reveal that technologies first appear in our scientific literature subset before reaching patents and social media, following traditional technology transfer patterns. While social media sentiment remains overwhelmingly bullish even during crypto winters, scientific and patent activity grows less tied to short-term sentiment, tracking overall market expansion in a virtuous cycle in which research precedes and enables economic growth that, in turn, funds further innovation. We release the DLT-Corpus and companion artifacts: LedgerBERT (+23% over BERT-base on DLT-specific Named Entity Recognition (NER) task), a sentiment analysis dataset of 23,301 crypto news headlines and descriptions, tools, and code.

Figures

Figures reproduced from arXiv: 2602.22045 by Jiahua Xu, Nikhil Vadgama, Paolo Tasca, Peter Devine, Walter Hernandez Cruz.

Figure 1
Figure 1. Figure 1: Overview of DLT-Corpus composition and patents 26.4k tokens. Token distribution: patents 43.5%, social media 37.6%, scientific literature 18.9% (Fig. 1a). 4.1.1 Scientific literature. Collection. We retrieved PDFs and metadata (authors, title, year, venue, references, licensing)12 from Semantic Scholar13 using domain￾specific queries (e.g., “Distributed Ledger Technology”, “Blockchain”, “Hashgraph”, “DAG”,… view at source ↗
Figure 2
Figure 2. Figure 2: Yearly growth of global cryptocurrency market capitaliza￾tion32and documents in the DLT￾Corpus. 2016 2018 2020 2022 Year 0.00 0.01 0.02 0.03 0.04 0.05 Proportion of Mentions Social Media Patents Scientific Lit. (a) Stablecoins 2016 2018 2020 2022 Year 0.000 0.005 0.010 0.015 0.020 0.025 0.030 0.035 Proportion of Mentions Social Media Patents Scientific Lit. (b) DEX 2014 2016 2018 2020 2022 Year 0.000 0.002… view at source ↗
Figure 4
Figure 4. Figure 4: Proportion of mentions per year for selected cryptocurrencies in the DLT-Corpus. The y-axis represents the relative frequency of [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Market sentiment in social media and yearly growth of [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

86 extracted references · 16 canonical work pages

  1. [1]

    Scientific publishing has a language problem.Nature Human Behaviour 2023 7:77, 7 (7 2023), 1019–1020

    2023. Scientific publishing has a language problem.Nature Human Behaviour 2023 7:77, 7 (7 2023), 1019–1020. doi:10.1038/s41562-023-01679-6

  2. [2]

    Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guil- herme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Pi- queres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von ...

  3. [3]

    Cohendet

    Fernand Amesse and P. Cohendet. 2001. Technology transfer revisited from the perspective of the knowledge-based economy.Research Policy30, 9 (12 2001), 1459–1478. doi:10.1016/S0048-7333(01)00162-7

  4. [4]

    Linara Axanova. 2012. U.S. Academic Technology Transfer Models: Traditional, Experimental And Hypothetical. (2012). http://lesnouvelles.lesi.org/lesnouvell es2012/lesnouvellesPDFJune2012/Axanova.pdf 44https://www.semanticscholar.org/ 45https://www.uspto.gov/terms-use-uspto-websites 46https://x.com/en/tos/previous/version_17 47https://x.com/en/tos/previous...

  5. [5]

    Nur Azmina, Mohamad Zamani, Jasy Liew, Suet Yan, and Ahmad Muhyiddin Yusof. 2022. XLNET-GRU Sentiment Regression Model for Cryptocurrency News in English and Malay. 36–42 pages. https://aclanthology.org/2022.fnp-1.5/

  6. [6]

    Adam Back, Matt Corallo, Luke Dashjr, Mark Friedenbach, Gregory Maxwell, Andrew Miller, Andrew Poelstra, Jorge Timón, and Pieter Wuille. 2014. Enabling Blockchain Innovations with Pegged Sidechains. (2014)

  7. [7]

    Anees Bahji, Laura Acion, Anne Marie Laslett, and Bryon Adinoff. 2023. Exclusion of the non-English-speaking world from the scientific literature: Recommenda- tions for change for addiction journals and publishers.Nordic Studies on Alcohol and Drugs40, 1 (2 2023), 6–13. doi:10.1177/14550725221102227

  8. [8]

    Leemon Baird and Atul Luykx. 2020. The Hashgraph Protocol: Efficient Asynchronous BFT for High-Throughput Distributed Ledgers.2020 Inter- national Conference on Omni-Layer Intelligent Systems, COINS 2020(8 2020). doi:10.1109/COINS49042.2020.9191430

  9. [9]

    Ballandies, Marcus M

    Mark C. Ballandies, Marcus M. Dapp, and Evangelos Pournaras. 2022. Decrypting distributed ledger design—taxonomy, classification and blockchain community evaluation.Cluster Computing25, 3 (6 2022), 1817–1838. doi:10.1007/S10586- 021-03256-W/FIGURES/12

  10. [10]

    Gruber, and Dirk Hovy

    Joachim Baumann, Paul Röttger, Aleksandra Urman, Albert Wendsjö, Flor Miriam Plaza-del Arco, Johannes B. Gruber, and Dirk Hovy. 2025. Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation. (9 2025). https://arxiv.org/abs/2509.08825v1

  11. [11]

    Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Mu- ralidharan, Yingyan Celine Lin, and Pavlo Molchanov. 2025. Small Language Models are the Future of Agentic AI. (6 2025). https://arxiv.org/abs/2506.02153v1

  12. [12]

    Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text.EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference(2019), 3615–3620. doi:10.18653/V1/D19-1371

  13. [13]

    Bruno Biais, Philip Bond, Jonathan Chiu, Rod Garratt, Niklas Haeusle, Shiyang Huang, Huisu Jang, Stephen Karolyi, Leonid Kogan, Jiasun Li, Tao Li, Evgeny Lyandres, Urban Jermann, Jonathan Payne, Julien Prat, Daniel Rabetti, Qihong Ruan, Fahad Saleh, Ville Savolainen, Donghwa Shin, Endong Yang, Jingjie Zhang, Shunming Zhang, Lin William Cong, Zhiheng He, a...

  14. [14]

    Blake Brittain. 2025. Judge explains order for New York Times in OpenAI copyright case | Reuters. https://www.reuters.com/legal/litigation/judge- explains-order-new-york-times-openai-copyright-case-2025-04-04/

  15. [15]

    Eric Budish. 2025. Trust at Scale: The Economic Limits of Cryptocurrencies and Blockchains.The Quarterly Journal of Economics140, 1 (1 2025), 1–62. doi:10.1093/QJE/QJAE033

  16. [16]

    Vitalik Buterin. 2014. Ethereum: A Next-Generation Smart Contract and Decen- tralized Application Platform. (2014). https://ethereum.org/content/whitepape r/whitepaper-pdf/Ethereum_Whitepaper_-_Buterin_2014.pdf

  17. [17]

    Richard Eckart De Castilho, Giulia Dore, Thomas Margoni, Penny Labropoulou, and Iryna Gurevych. 2018. A Legal Perspective on Training Models for Natural Language Processing. https://aclanthology.org/L18-1202/

  18. [18]

    Chang and Benjamin K

    Tyler A. Chang and Benjamin K. Bergen. 2022. Word Acquisition in Neural Language Models.Transactions of the Association for Computational Linguistics 10 (1 2022), 1–16. https://aclanthology.org/2022.tacl-1.1/

  19. [19]

    E. Chen, N. Roche, Y.-H. Tseng, W. Hernandez, J. Shangguan, and A. Moore. 2023. Conversion of Legal Agreements into Smart Legal Contracts using NLP. InACM Web Conference 2023 - Companion of the World Wide Web Conference, WWW

  20. [20]

    Lin William Cong, Yuanyu Qu, and Guojun Wang. 2025. Blockchains for envi- ronmental monitoring: theory and empirical evidence from China.Review of Finance29, 5 (9 2025), 1303–1336. doi:10.1093/ROF/RFAF033

  21. [21]

    Davidson, Darja Wischerath, Daniel Racek, Douglas A

    Brittany I. Davidson, Darja Wischerath, Daniel Racek, Douglas A. Parry, Emily Godwin, Joanne Hinds, Dirk van der Linden, Jonathan F. Roscoe, Laura Ayra- vainen, and Alicia G. Cork. 2023. Platform-controlled social media APIs threaten open science.Nature Human Behaviour 2023 7:127, 12 (11 2023), 2054–2057. doi:10.1038/s41562-023-01750-2

  22. [22]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova Google, and A I Language. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.Proceedings of the 2019 Conference of the North(2019), 4171–4186. doi:10.18653/V1/N19-1423

  23. [23]

    Wenzhi Ding, Chen Lin, Yichen Luo, and Jiahua Xu. 2025. Decompose Market Manipulation Strategies: Evidence from On-chain Meme Coin Market. (9 2025). doi:10.2139/SSRN.5953738

  24. [24]

    Honglin Fu, Yebo Feng, Cong Wu, and Jiahua Xu. 2025. \textsc{Perseus}: Tracing the Masterminds Behind Cryptocurrency Pump-and-Dump Schemes. (3 2025). https://arxiv.org/abs/2503.01686v1

  25. [25]

    Yu Gai, Liyi Zhou, Kaihua Qin, Dawn Song, and Arthur Gervais. 2023. Blockchain Large Language Models. (4 2023). https://arxiv.org/abs/2304.12749v2

  26. [26]

    Amish Garg, Tanav Shah, Vinay Kumar Jain, and Raksha Sharma. 2021. Cryp- Top12: A Dataset for Cryptocurrency Price Movement Prediction from Tweets and Historical Prices.Proceedings - 20th IEEE International Conference on Machine Learning and Applications, ICMLA 2021(2021), 379–384. doi:10.1109/ICMLA529 53.2021.00065

  27. [27]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM64, 12 (11 2021), 86–92. doi:10.1145/3458723

  28. [28]

    Anjee Gorkhali, Ling Li, and Asim Shrestha. 2020. Blockchain: a literature review. Journal of Management Analytics7, 3 (7 2020), 321–343. doi:10.1080/23270012.2 020.1801529

  29. [29]

    David Grangier, Angelos Katharopoulos, Pierre Ablin, and Awni Hannun Apple

  30. [30]

    Dominique Guégan and Thomas Renault. 2021. Does investor sentiment on social media provide robust information for Bitcoin returns predictability?Finance Research Letters38 (1 2021), 101494. doi:10.1016/J.FRL.2020.101494

  31. [31]

    Vincent Gurgul, Stefan Lessmann, and Wolfgang Karl Härdle. 2025. Deep learning and NLP in cryptocurrency forecasting: Integrating financial, blockchain, and social media data.International Journal of Forecasting(3 2025). doi:10.1016/J.IJ FORECAST.2025.02.007

  32. [32]

    Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks.Proceedings of the Annual Meeting of the Associa- tion for Computational Linguistics(2020), 8342–8360. doi:10.18653/V1/2020.ACL- MAIN.740

  33. [33]

    Eric Harris-Braun, Arthur Brock, and Paul D’aoust. [n. d.]. Holochain Distributed Coordination by Scaled Consent, not Global Consensus. ([n. d.]). doi:10.1145/32 2186.322188

  34. [34]

    Melissa Heikkila, Chris Cook, and Clara Murray. 2025. America’s top companies keep talking about AI — but can’t explain the upsides. https://www.ft.com/con tent/e93e56df-dd9b-40c1-b77a-dba1ca01e473

  35. [35]

    Walter Hernandez Cruz, Firas Dahi, Yebo Feng, Jiahua Xu, Aanchal Malhotra, and Paolo Tasca. 2025. AMM-based DEX on the XRP Ledger.2025 IEEE International Conference on Blockchain and Cryptocurrency (ICBC)(6 2025), 1–10. doi:10.1109/ ICBC64466.2025.11114626

  36. [36]

    Walter Hernandez Cruz, Kamil Tylinski, Alastair Moore, Niall Roche, Nikhil Vadgama, Horst Treiblmaier, Jiangbo Shangguan, Paolo Tasca, and Jiahua Xu

  37. [37]

    Walter Hernandez Cruz, Jiahua Xu, Paolo Tasca, and Carlo Campajola. 2024. No Questions Asked: Effects of Transparency on Stablecoin Liquidity During the Collapse of Silicon Valley Bank. (2024). https://arxiv.org/abs/2407.11716

  38. [38]

    Kia Jahanbin, Mohammad Ali Zare Chahooki, and Fereshte Rahmanian. 2023. Database of influencers’ tweets in cryptocurrency (2021-2023). 2 (2023). doi:10.1 7632/8FBDHH72GS.2

  39. [39]

    Lily Jamali. 2025. AI firm Anthropic agrees to pay authors $1.5bn for pirating work - BBC News. https://www.bbc.co.uk/news/articles/c5y4jpg922qo

  40. [40]

    Armand Joulin, Édouard Grave, Piotr Bojanowski, and Tomáš Mikolov. 2017. Bag of Tricks for Efficient Text Classification. 427–431 pages. https://aclanthology.o rg/E17-2068/

  41. [41]

    Martin Juan, José Bucher, and Marco Martini. 2024. Fine-Tuned ’Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classifi- cation. (6 2024). https://arxiv.org/abs/2406.08660v2

  42. [42]

    Inwon Kang, Maruf Ahmed Mridul, Abraham Sanders, Yao Ma, Thilanka Mu- nasinghe, Aparna Gupta, and Oshani Seneviratne. 2024. Deciphering Crypto Twitter.Proceedings of the 16th ACM Web Science Conference, WebSci 2024(5 2024), 331–342. doi:10.1145/3614419.3644026

  43. [43]

    Jaehyun Kim, Thi Thu Huong Le, Sangmyeong Lee, and Howon Kim. 2024. Ethereum Smart Contracts Vulnerabilities Detection Leveraging Fine-Tuning DistilBERT.International Conference on Platform Technology and Service(2024), 133–138. doi:10.1109/PLATCON63925.2024.10830749

  44. [44]

    Kate Knibbs. 2025. Meta Secretly Trained Its AI on a Notorious Piracy Database, Newly Unredacted Court Docs Reveal | WIRED. https://www.wired.com/story/ new-documents-unredacted-meta-copyright-ai-lawsuit/

  45. [45]

    Olivier Kraaijeveld and Johannes De Smedt. 2020. The predictive power of public Twitter sentiment for forecasting cryptocurrency prices.Journal of International Financial Markets, Institutions and Money65 (3 2020), 101188. doi:10.1016/J.INTF IN.2020.101188

  46. [46]

    Xianzhi Li, Samuel Chan, Xiaodan Zhu, Yulong Pei, Zhiqiang Ma, Xiaomo Liu, and Sameena Shah. 2023. Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks.EMNLP 2023 - 2023 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Industry Track(2023), 408–422. doi:10.18653/V1/2...

  47. [47]

    Yuan Li, Bingqiao Luo, Qian Wang, Nuo Chen, Xu Liu, and Bingsheng He. 2024. CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading.EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference(2024), 1094–1106. doi:10.18653/V1/20 24.EMNLP-MAIN.63 Hernandez Cruz et al

  48. [48]

    Zichao Li. 2025. Knowledge-Grounded Detection of Cryptocurrency Scams with Retrieval-Augmented LMs. (8 2025), 40–48. doi:10.18653/V1/2025.KNOWLLM-1.4

  49. [49]

    Gordon Y Liao and John Caramichael. 2022. Stablecoins: Growth Potential and Impact on Banking.International Finance Discussion Paper2022, 1334 (2022), 1–26. doi:10.17016/ifdp.2022.1334

  50. [50]

    Lo and Francesca Medda

    Yuen C. Lo and Francesca Medda. 2020. Assets on the blockchain: An empirical study of Tokenomics.Information Economics and Policy53 (12 2020), 100881. doi:10.1016/J.INFOECOPOL.2020.100881

  51. [51]

    Jinghui Lu, Maeve Henchion, and Brian Mac Namee. 2020. Diverging Divergences: Examining Variants of Jensen Shannon Divergence for Corpus Comparison Tasks. (2020), 11–16

  52. [52]

    Yichen Luo, Yebo Feng, Jiahua Xu, and Yang Liu. 2026. Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of- Thought Reasoning.Proceedings of the ACM Web Conference 2026 (WWW ’26), April 13â•fi17, 2026, Dubai, United Arab Emirates1 (1 2026). doi:10.1145/3774904. 3792635

  53. [53]

    Sean McNally, Jason Roche, and Simon Caton. 2018. Predicting the Price of Bitcoin Using Machine Learning.International Euromicro Conference on Parallel, Distributed and Network-Based Processing(6 2018), 339–343. doi:10.1109/PDP201 8.2018.00060

  54. [54]

    Roberto Moncada, Enrico Ferro, Maurizio Fiaschetti, and Francesca Medda. 2024. Blockchain Tokens, Price Volatility, and Active User Base: An Empirical Analysis Based on Tokenomics.International Journal of Financial Studies 2024, Vol. 12, Page 10712, 4 (10 2024), 107. doi:10.3390/IJFS12040107

  55. [55]

    Satoshi Nakamoto. 2008. Bitcoin: A peer-to-Peer Electronic Cash System. 552– 557 pages. https://bitcoin.org/bitcoin.pdf

  56. [56]

    Leonardo Nizzoli, Serena Tardelli, Marco Avvenuti, Stefano Cresci, Maurizio Tesconi, and Emilio Ferrara. 2020. Charting the Landscape of Online Cryptocur- rency Manipulation.IEEE Access8 (2020), 113230–113245. doi:10.1109/ACCESS .2020.3003370

  57. [57]

    Anna Paula Pawlicka Maule and Kristen Marie Johnson. 2021. Cryptocurrency Day Trading and Framing Prediction in Microblog Discourse.Proceedings of the 3rd Workshop on Economics and Natural Language Processing, ECONLP 2021 (2021), 82–92. doi:10.18653/V1/2021.ECONLP-1.11

  58. [58]

    Branislav Pecher, Ivan Srba, and Maria Bielikova. 2025. Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance. (11 2025), 165–184. doi:10.18653 /V1/2025.EMNLP-MAIN.9

  59. [59]

    Pannier, Ebtesam Almazrouei, and Julien Launay

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra-Aimée Cojo- caru, Hamza Alobeidli, Alessandro Cappelli, B. Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperform- ing Curated Corpora with Web Data Only.Neural Information Processing Systems (2023)

  60. [60]

    Arif Perdana, Alastair Robb, Vivek Balachandran, and Fiona Rohde. 2021. Dis- tributed ledger technology: Its evolutionary path and the road ahead.Information & Management58, 3 (4 2021), 103316. doi:10.1016/J.IM.2020.103316

  61. [61]

    Audrey Pope. 2024. NYT v. OpenAI: The Times’s About-Face - Harvard Law Review. https://harvardlawreview.org/blog/2024/04/nyt-v-openai-the-timess- about-face/

  62. [62]

    Jacob Portes, Alex Trott, Sam Havens, Daniel King, Abhinav Venigalla, Moin Nadeem, Nikhil Sardana, Daya Khudia, and Jonathan Frankle. 2023. MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining.Advances in Neural Information Processing Systems36 (12 2023). https://arxiv.org/abs/2312.17482v2

  63. [63]

    Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J

    Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of machine learning research(2019)

  64. [64]

    Mayank Raikwar, Nikita Polyanskii, and Sebastian Muller. 2024. SoK: DAG- based Consensus Protocols.2024 IEEE International Conference on Blockchain and Cryptocurrency, ICBC 2024(2024). doi:10.1109/ICBC59979.2024.10634358

  65. [65]

    Pornpanit Rasivisuth, Maurizio Fiaschetti, and Francesca Medda. 2024. An investigation of sentiment analysis of information disclosure during Initial Coin Offering (ICO) on the token return.International Review of Financial Analysis95 (10 2024), 103437. doi:10.1016/J.IRFA.2024.103437

  66. [66]

    R. L. Rivest, A. Shamir, and L. Adleman. 1978. A method for obtaining digital signatures and public-key cryptosystems.Commun. ACM21, 2 (2 1978), 120–126. doi:10.1145/359340.359342

  67. [67]

    Sougata Sarkar, Aditya Badwal, Amartya Roy, Koustav Rudra, and Kripabandhu Ghosh. 2025. CryptOpiQA: A new Opinion and Question Answering dataset on Cryptocurrency. 11107–11120 pages. doi:10.5281/zenodo.14469000

  68. [68]

    Pavlo Seroyizhko, Zhanel Zhexenova, Muhammad Zohaib Shafiq, Fabio Merizzi, Andrea Galassi, and Federico Ruggeri. 2022. A Sentiment and Emotion Annotated Dataset for Bitcoin Price Forecasting Based on Reddit Posts.FinNLP 2022 - 4th Workshop on Financial Technology and Natural Language Processing, Proceedings of the Workshop(2022), 203–210. doi:10.18653/V1/...

  69. [69]

    Rion Snow, Brendan O’connor, Daniel Jurafsky, and Andrew Y Ng. 2008. Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. 254–263 pages. https://aclanthology.org/D08-1027/

  70. [70]

    Pollard, Eric Lehman, Alistair E

    Thomas Sounack, Joshua Davis, Brigitte Durieux, Antoine Chaffin, Tom J. Pollard, Eric Lehman, Alistair E. W. Johnson, Matthew McDermott, Tristan Naumann, and Charlotta Lindvall. 2025. BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP. (6 2025). https: //arxiv.org/abs/2506.10896v1

  71. [71]

    Jonathan Stempel. 2025. SEC ends lawsuit against Ripple, company to pay $125 million fine | Reuters. https://www.reuters.com/legal/government/sec-ends- lawsuit-against-ripple-company-pay-125-million-fine-2025-08-08/

  72. [72]

    Jianguo Sun, Yifan Jia, Yanbin Wang, Ye Tian, and Sheng Zhang. 2025. Ethereum fraud detection via joint transaction language model and graph representation learning.Information Fusion120 (8 2025), 103074. doi:10.1016/J.INFFUS.2025.10 3074

  73. [73]

    Paolo Tasca and Claudio J. Tessone. 2017. Taxonomy of Blockchain Technologies. Principles of Identification and Classification.Ledger4 (5 2017), 1–39. doi:10.5 195/LEDGER.2019.140

  74. [74]

    Benjamin Warner, Antoine Chaffin,†Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inferen...

  75. [75]

    Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Ap- pleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan Willem Boiten, Luiz Bonino da Silva Santos, Philip E

    Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Ap- pleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Mercè Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra Gonzalez-Beltran,...

  76. [76]

    Gavin Wood. 2016. Polkadot: Vision for a Heterogeneous Multi-Chain Frame- work. (2016)

  77. [77]

    Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos, Ali Farhadi, and Ludwig Schmidt. 2023. Stable and low-precision training for large-scale vision-language models. InProceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA

  78. [78]

    Yang Xiao, Mohan Jiang, Jie Sun, Keyu Li, Jifan Lin, Yumin Zhuang, Ji Zeng, Shijie Xia, Qishuo Hua, Xuefeng Li, Xiaojie Cai, Tongyu Wang, Yue Zhang, Liming Liu, Xia Wu, Jinlong Hou, Yuan Cheng, Wenjie Li, Xiang Wang, Dequan Wang, and Pengfei Liu. 2025. LIMI: Less is More for Agency. (9 2025). https: //arxiv.org/abs/2509.17567v2

  79. [79]

    Yong Xie, Karan Aggarwal, and Aitzaz Ahmad. 2024. Efficient Continual Pre- training for Building Domain Specific Large Language Models.Findings of the Association for Computational Linguistics ACL 2024(2024), 10184–10201. doi:10.18653/V1/2024.FINDINGS-ACL.606

  80. [80]

    Jiahua Xu, Krzysztof Paruch, Simon Cousaert, and Yebo Feng. 2023. SoK: De- centralized Exchanges (DEX) with Automated Market Maker (AMM) Protocols. Comput. Surveys55, 11 (11 2023), 1–50. doi:10.1145/3570639

Showing first 80 references.