REVIEW 4 major objections 5 minor 86 references
A 2.98-billion-token corpus spanning science, patents, and social media shows that DLT concepts first appear in research before markets.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-02 20:46 UTC pith:AGLSTIY6
load-bearing objection The corpus is a real resource worth having; the diffusion and market-lead analyses overreach, but that's fixable. the 4 major comments →
DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that DLT-Corpus, the largest domain-specific text collection for distributed ledger technology, enables both specialized NLP and cross-domain studies of innovation. Using it, the authors find that three economically significant technologies (stablecoins, DEXs, AMMs) consistently appear in scientific literature before they appear in patents or social media, matching a traditional technology-transfer pattern. They also find an asymmetric lagged correlation: scientific publications lead market capitalization by two years (rho=0.95, p<0.001) and the correlation decays when the market leads, while patents show a symmetric correlation and social media follows the market. The c
What carries the argument
The central object is the corpus itself, which aggregates three timestamped document streams: scientific literature (37,440 open-access publications), US patents (49,023 records), and social media posts (22 million). The analysis relies on the corpus's temporal metadata to compute lagged correlations and first-mention dates, and on its keyword density (8.7 times higher than general web corpora) to argue it provides concentrated learning signal for domain-adapted models. The filtering pipeline, including a BERT-based domain relevance filter trained on an existing DLT NER dataset and manual pruning, defines the corpus's membership and thus drives the first-mention ordering.
Load-bearing premise
The technology-transfer ordering rests on first-mention dates derived from a corpus whose scientific subset was filtered by a BERT-based relevance model and whose patents were selected by a simple keyword search; different query or threshold choices could change the order.
What would settle it
For any of the three analyzed technologies (stablecoins, DEXs, AMMs), find a patent or social media post dated earlier than the corpus's first scientific mention of that term; such a document would break the claimed science-first ordering.
If this is right
- Tracking scientific literature can serve as an early-warning system for emerging DLT technologies, since concepts first appear there.
- Scientific output appears to lead cryptocurrency market growth by about two years, suggesting research is a leading indicator for DLT market expansion.
- Domain-adapted language models trained on DLT-Corpus outperform general models on DLT-specific NER, showing concentrated terminology matters.
- The corpus enables integrated analysis across scientific, patent, and social discourse that previously required assembling separate, often inaccessible datasets.
- Because social media coverage stops at 2023, the corpus is a fixed snapshot and cannot track post-2023 community discourse.
Where Pith is reading between the lines
- The first-mention ordering (science before patents before social media) is sensitive to how each subset was collected and filtered; a differently filtered corpus could yield a different diffusion ordering, so the technology-transfer conclusion should be tested on alternative query and filtering choices.
- The two-year lead of scientific publications over market cap is based on annual aggregates; finer-grained data could reveal shorter lags or lead-lag dynamics within quarters, and the correlation alone does not establish causation.
- The persistently bullish social sentiment may reflect the cryptocurrency-focused composition of the social media sources and the sentiment labeling method rather than an objective market mood; a sample of general social media would be a useful check.
- The corpus excludes news articles for copyright reasons; a news stream, if added, might show that financial journalism acts as an intermediary between research and market sentiment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DLT-Corpus, a large multi-source text collection for the Distributed Ledger Technology domain, containing 2.98 billion tokens from 37,440 scientific publications, 49,023 USPTO patents, and 22.03 million Twitter posts. The authors describe the construction and quality assessment of the corpus, train a domain-adapted language model (LedgerBERT), release a crowdsourced sentiment-analysis dataset, and present two utility analyses: (i) correlations between document volumes, market capitalization, and sentiment; and (ii) technology diffusion across scientific literature, patents, and social media. The headline claims are that technologies originate in scientific literature before reaching patents and social media, and that scientific publications lead market expansion by two years (ρ=0.95, p<0.001). The resource is publicly released with code, models, and datasheet-style documentation.
Significance. If the claims hold, DLT-Corpus is a substantial and valuable contribution: it is described as the largest domain-specific DLT text collection, it is publicly released, and it ships reproducible artifacts (code, models, datasets) as well as a FAIR-aligned datasheet. The quality assessment against general-purpose corpora is a useful benchmark. The temporal and cross-source analyses are of interest to innovation-diffusion and bibliometric communities. However, the two central analytical claims — technology-transfer ordering and science-leading-market causality — are currently supported only by correlations and normalized mention proportions, and the evaluation of LedgerBERT is partly in-domain by construction. The resource itself is likely useful even if those analytical claims are weakened; the paper's contribution should be reframed accordingly.
major comments (4)
- [§6.1, Figs. 3–4] The central claim that technologies 'originate in scientific literature before reaching patents and social media' is not established by the reported evidence. Figures 3 and 4 show normalized yearly proportions, not first-mention dates per technology or per document, and no statistical test for ordering is provided. More importantly, the three subsets have different temporal coverage (scientific: 1978–2025; patents: 1990–2025; social media: 2013–mid 2023; §4.1.3) and different collection filters. The patents subset uses only the search terms 'Distributed Ledger Technology' and 'blockchain' (§4.1.2), which can miss early patents using terms such as 'cryptocurrency', 'digital asset', or 'smart contract', biasing patent first-mentions late. The social media subsets are aggregated from Kaggle and academic datasets whose earliest coverage and sampling gaps are not documented; if those datasets
- [Abstract; §6, Tables 4–5] The abstract states that 'scientific publications lead market expansion by two years (ρ=0.95, p<0.001)', but Table 5 shows that social media also exhibits ρ=0.95 at lag −3, and patents exhibit ρ=0.95 at lag −2 and ρ=0.93 at lag −3. The claim that scientific literature uniquely or particularly leads is not supported by any test for differences among lags or among document types. Moreover, the correlations are computed on annual data from 2013–2024 (n=12 for scientific literature and patents, n=11 for social media) with no correction for autocorrelation or multiple testing; Spearman's ρ on such short, highly autocorrelated series can be large and significant even when no true lead–lag relationship exists. The full text also reports concurrent ρ=0.76 for scientific literature (Table 4), so the 'two-year lead' claim rests on a single lagged correlation without a formal comparison or robustne
- [§4.1.1 and §5.1] The primary evaluation of LedgerBERT is partly circular. The scientific-literature subset of DLT-Corpus was filtered by fine-tuning BERT-base-cased on the NER dataset from [36] and predicting domain entities (§4.1.1). LedgerBERT is then evaluated on exactly that same NER dataset from [36] as the 'primary evaluation' (§5.1). This means the model is trained on a corpus that was selected using labels from the very evaluation set, and the reported +23% over BERT-base and +3.5% over SciBERT may be inflated by construction. The authors note that the dataset 'derives from scientific literature, matching our corpus composition', but they do not discuss the selection overlap. An independent held-out evaluation (e.g., NER on patents or social media, or a DLT NER dataset not used in corpus filtering) is needed to support the claim that LedgerBERT improves DLT-specific NER because of domain-adaptive
- [§9 Limitations] The limitations section is incomplete regarding the two headline analyses. It states that 'marginally relevant DLT papers may remain in the dataset' but does not acknowledge that the science-first diffusion order and the market-lead correlations could be artifacts of (i) the patents subset's narrow keyword selection, (ii) undocumented temporal coverage or gaps in the social-media aggregations, or (iii) the scientific-filtering model's vocabulary biases. Because these threats bear directly on the paper's central claims, they should be addressed explicitly, preferably with supplementary analyses that restrict all three subsets to a common overlapping period and/or use alternative keyword vocabularies.
minor comments (5)
- [§6.1 text] The text refers to 'Fig. 2a' for stablecoins, 'Fig. 2b' for DEXs, and 'Fig. 2c' for AMMs, but the figures are numbered Fig. 3(a)–(c). Please correct the cross-references.
- [§6.2] The sentence 'Comparing Fig. 1d with Fig. 5' appears to refer to Fig. 1(d), the temporal evolution panel, but the connection between sentiment and document growth is asserted rather than quantified. Consider reporting a correlation between sentiment fractions and document volumes, or clearly labeling this as a visual comparison.
- [§4.2] The description of the sentiment-label construction would benefit from more detail: the 'median minimum votes' filter and the 25th/75th percentile boundaries are mentioned, but the exact thresholds and the number of examples excluded at each step are not reported. This is a reproducibility concern for the sentiment dataset.
- [Appendix A, Table 8] The social-media subset is described as 'Twitter/X posts', but several upstream sources are Kaggle datasets (e.g., 'bitcoin-tweets', 'crypto-tweets'). It would be useful to document which sources cover which time ranges, since this directly affects the diffusion analysis.
- [Throughout] Minor typographical issues: 'cryptocurrencies price prediction' in the abstract; 'T&Cs at collection time' should be 'terms and conditions'; the footnote on line 30 of §5.2 is incomplete. A careful proofread is recommended.
Circularity Check
LedgerBERT's flagship NER gain is measured on the same [36] annotations used to filter its own pretraining corpus; the rest of the claims are largely self-contained.
specific steps
-
other
[§4.1.1 (Domain filtering) and §5.1 (Primary evaluation: in-domain NER), Table 2]
"To ensure relevance, we fine-tuned BERT-base-cased on the NER dataset from [36] and predicted domain-specific entities in each document. // NER serves as the primary evaluation of corpus quality because performance directly reflects how well the model learned domain-specific terminology. We use the DLT-focused NER dataset from [36]... This dataset derives from scientific literature, matching our corpus composition."
The same authors' [36] NER annotations are load-bearing twice: they train the BERT filter that selects which scientific documents enter DLT-Corpus (§4.1.1), and they are the benchmark on which LedgerBERT is evaluated (§5.1) after continued pretraining on that filtered corpus. The reported +23.05% NER gain over BERT-base is therefore measured on a label scheme that already shaped the training corpus: documents were retained only if a model trained on [36] entities scored them highly. The evaluation is not a mathematical identity (the model must still learn the mapping), but as "the primary evaluation of corpus quality" it partly measures the same signal used in construction, so the utility claim is partly self-referential rather than independently validated.
full rationale
The corpus itself, its size, its public release, and the external keyword-density comparison (Table 1) are not defined in terms of the paper's conclusions, so the core dataset contribution is self-contained. The innovation-diffusion and lagged-correlation analyses are empirical exercises on the released artifacts; they may be confounded by unequal temporal coverage and keyword-based patent selection (a correctness risk the authors' §9 does not address), but no equation reduces the observed ordering to the paper's inputs. The one clear circularity is the [36] filter/evaluation loop: LedgerBERT is pretrained on a scientific subset selected by a BERT model fine-tuned on [36]'s NER labels, and then the paper's headline utility result is F1 on that same [36] NER dataset. Because [36] is the authors' own prior work and is invoked both for corpus construction and as the primary benchmark, the evaluation is in-domain by construction. Yet the model's gain is small and not logically forced, and the out-of-domain sentiment benchmark and external corpus comparisons provide independent content, so the appropriate score is 4 rather than a higher one.
Axiom & Free-Parameter Ledger
free parameters (3)
- Domain-filtering BERT thresholds =
max prediction score > 0.995, or median score at/above subset-wide median
- Sentiment label percentile boundaries =
25th and 75th percentiles of normalized vote percentages
- Manual removal of 570 papers =
570 papers removed
axioms (4)
- domain assumption Semantic Scholar open-access papers with the DLT queries form a representative sample of DLT scientific literature.
- domain assumption Keyword-search patents with 'Distributed Ledger Technology'/'blockchain' capture the relevant patent landscape.
- domain assumption Aggregated Kaggle/industry tweet collections are a representative snapshot of DLT community discourse.
- domain assumption First-mention dates in the corpus reflect when technologies actually first appeared in each community.
read the original abstract
We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents spanning scientific literature (37,440 publications), United States Patent and Trademark Office (USPTO) patents (49,023 filings), and social media (22 million posts). Existing Natural Language Processing (NLP) resources for DLT focus narrowly on cryptocurrency price prediction and smart contracts, leaving domain-specific language underexplored despite the sector's ~$3 trillion market capitalization and rapid technological evolution. We demonstrate DLT-Corpus' utility by analyzing patterns of technology emergence and market-innovation correlations. Findings reveal that technologies first appear in our scientific literature subset before reaching patents and social media, following traditional technology transfer patterns. While social media sentiment remains overwhelmingly bullish even during crypto winters, scientific and patent activity grows less tied to short-term sentiment, tracking overall market expansion in a virtuous cycle in which research precedes and enables economic growth that, in turn, funds further innovation. We release the DLT-Corpus and companion artifacts: LedgerBERT (+23% over BERT-base on DLT-specific Named Entity Recognition (NER) task), a sentiment analysis dataset of 23,301 crypto news headlines and descriptions, tools, and code.
Figures
Reference graph
Works this paper leans on
-
[1]
Scientific publishing has a language problem.Nature Human Behaviour 2023 7:77, 7 (7 2023), 1019–1020
2023. Scientific publishing has a language problem.Nature Human Behaviour 2023 7:77, 7 (7 2023), 1019–1020. doi:10.1038/s41562-023-01679-6
-
[2]
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guil- herme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Pi- queres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von ...
Pith/arXiv arXiv 2025
-
[3]
Fernand Amesse and P. Cohendet. 2001. Technology transfer revisited from the perspective of the knowledge-based economy.Research Policy30, 9 (12 2001), 1459–1478. doi:10.1016/S0048-7333(01)00162-7
-
[4]
Linara Axanova. 2012. U.S. Academic Technology Transfer Models: Traditional, Experimental And Hypothetical. (2012). http://lesnouvelles.lesi.org/lesnouvell es2012/lesnouvellesPDFJune2012/Axanova.pdf 44https://www.semanticscholar.org/ 45https://www.uspto.gov/terms-use-uspto-websites 46https://x.com/en/tos/previous/version_17 47https://x.com/en/tos/previous...
2012
-
[5]
Nur Azmina, Mohamad Zamani, Jasy Liew, Suet Yan, and Ahmad Muhyiddin Yusof. 2022. XLNET-GRU Sentiment Regression Model for Cryptocurrency News in English and Malay. 36–42 pages. https://aclanthology.org/2022.fnp-1.5/
2022
-
[6]
Adam Back, Matt Corallo, Luke Dashjr, Mark Friedenbach, Gregory Maxwell, Andrew Miller, Andrew Poelstra, Jorge Timón, and Pieter Wuille. 2014. Enabling Blockchain Innovations with Pegged Sidechains. (2014)
2014
-
[7]
Anees Bahji, Laura Acion, Anne Marie Laslett, and Bryon Adinoff. 2023. Exclusion of the non-English-speaking world from the scientific literature: Recommenda- tions for change for addiction journals and publishers.Nordic Studies on Alcohol and Drugs40, 1 (2 2023), 6–13. doi:10.1177/14550725221102227
-
[8]
Leemon Baird and Atul Luykx. 2020. The Hashgraph Protocol: Efficient Asynchronous BFT for High-Throughput Distributed Ledgers.2020 Inter- national Conference on Omni-Layer Intelligent Systems, COINS 2020(8 2020). doi:10.1109/COINS49042.2020.9191430
arXiv 2020
-
[9]
Mark C. Ballandies, Marcus M. Dapp, and Evangelos Pournaras. 2022. Decrypting distributed ledger design—taxonomy, classification and blockchain community evaluation.Cluster Computing25, 3 (6 2022), 1817–1838. doi:10.1007/S10586- 021-03256-W/FIGURES/12
-
[10]
Joachim Baumann, Paul Röttger, Aleksandra Urman, Albert Wendsjö, Flor Miriam Plaza-del Arco, Johannes B. Gruber, and Dirk Hovy. 2025. Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation. (9 2025). https://arxiv.org/abs/2509.08825v1
arXiv 2025
-
[11]
Peter Belcak, Greg Heinrich, Shizhe Diao, Yonggan Fu, Xin Dong, Saurav Mu- ralidharan, Yingyan Celine Lin, and Pavlo Molchanov. 2025. Small Language Models are the Future of Agentic AI. (6 2025). https://arxiv.org/abs/2506.02153v1
Pith/arXiv arXiv 2025
-
[12]
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. SciBERT: A Pretrained Language Model for Scientific Text.EMNLP-IJCNLP 2019 - 2019 Conference on Empirical Methods in Natural Language Processing and 9th International Joint Conference on Natural Language Processing, Proceedings of the Conference(2019), 3615–3620. doi:10.18653/V1/D19-1371
-
[13]
Bruno Biais, Philip Bond, Jonathan Chiu, Rod Garratt, Niklas Haeusle, Shiyang Huang, Huisu Jang, Stephen Karolyi, Leonid Kogan, Jiasun Li, Tao Li, Evgeny Lyandres, Urban Jermann, Jonathan Payne, Julien Prat, Daniel Rabetti, Qihong Ruan, Fahad Saleh, Ville Savolainen, Donghwa Shin, Endong Yang, Jingjie Zhang, Shunming Zhang, Lin William Cong, Zhiheng He, a...
-
[14]
Blake Brittain. 2025. Judge explains order for New York Times in OpenAI copyright case | Reuters. https://www.reuters.com/legal/litigation/judge- explains-order-new-york-times-openai-copyright-case-2025-04-04/
2025
-
[15]
Eric Budish. 2025. Trust at Scale: The Economic Limits of Cryptocurrencies and Blockchains.The Quarterly Journal of Economics140, 1 (1 2025), 1–62. doi:10.1093/QJE/QJAE033
-
[16]
Vitalik Buterin. 2014. Ethereum: A Next-Generation Smart Contract and Decen- tralized Application Platform. (2014). https://ethereum.org/content/whitepape r/whitepaper-pdf/Ethereum_Whitepaper_-_Buterin_2014.pdf
2014
-
[17]
Richard Eckart De Castilho, Giulia Dore, Thomas Margoni, Penny Labropoulou, and Iryna Gurevych. 2018. A Legal Perspective on Training Models for Natural Language Processing. https://aclanthology.org/L18-1202/
2018
-
[18]
Chang and Benjamin K
Tyler A. Chang and Benjamin K. Bergen. 2022. Word Acquisition in Neural Language Models.Transactions of the Association for Computational Linguistics 10 (1 2022), 1–16. https://aclanthology.org/2022.tacl-1.1/
2022
-
[19]
E. Chen, N. Roche, Y.-H. Tseng, W. Hernandez, J. Shangguan, and A. Moore. 2023. Conversion of Legal Agreements into Smart Legal Contracts using NLP. InACM Web Conference 2023 - Companion of the World Wide Web Conference, WWW
2023
-
[20]
Lin William Cong, Yuanyu Qu, and Guojun Wang. 2025. Blockchains for envi- ronmental monitoring: theory and empirical evidence from China.Review of Finance29, 5 (9 2025), 1303–1336. doi:10.1093/ROF/RFAF033
-
[21]
Davidson, Darja Wischerath, Daniel Racek, Douglas A
Brittany I. Davidson, Darja Wischerath, Daniel Racek, Douglas A. Parry, Emily Godwin, Joanne Hinds, Dirk van der Linden, Jonathan F. Roscoe, Laura Ayra- vainen, and Alicia G. Cork. 2023. Platform-controlled social media APIs threaten open science.Nature Human Behaviour 2023 7:127, 12 (11 2023), 2054–2057. doi:10.1038/s41562-023-01750-2
-
[22]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova Google, and A I Language. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.Proceedings of the 2019 Conference of the North(2019), 4171–4186. doi:10.18653/V1/N19-1423
-
[23]
Wenzhi Ding, Chen Lin, Yichen Luo, and Jiahua Xu. 2025. Decompose Market Manipulation Strategies: Evidence from On-chain Meme Coin Market. (9 2025). doi:10.2139/SSRN.5953738
-
[24]
Honglin Fu, Yebo Feng, Cong Wu, and Jiahua Xu. 2025. \textsc{Perseus}: Tracing the Masterminds Behind Cryptocurrency Pump-and-Dump Schemes. (3 2025). https://arxiv.org/abs/2503.01686v1
Pith/arXiv arXiv 2025
-
[25]
Yu Gai, Liyi Zhou, Kaihua Qin, Dawn Song, and Arthur Gervais. 2023. Blockchain Large Language Models. (4 2023). https://arxiv.org/abs/2304.12749v2
Pith/arXiv arXiv 2023
-
[26]
Amish Garg, Tanav Shah, Vinay Kumar Jain, and Raksha Sharma. 2021. Cryp- Top12: A Dataset for Cryptocurrency Price Movement Prediction from Tweets and Historical Prices.Proceedings - 20th IEEE International Conference on Machine Learning and Applications, ICMLA 2021(2021), 379–384. doi:10.1109/ICMLA529 53.2021.00065
arXiv 2021
-
[27]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé Iii, and Kate Crawford. 2021. Datasheets for datasets. Commun. ACM64, 12 (11 2021), 86–92. doi:10.1145/3458723
doi:10.1145/3458723 2021
-
[28]
Anjee Gorkhali, Ling Li, and Asim Shrestha. 2020. Blockchain: a literature review. Journal of Management Analytics7, 3 (7 2020), 321–343. doi:10.1080/23270012.2 020.1801529
-
[29]
David Grangier, Angelos Katharopoulos, Pierre Ablin, and Awni Hannun Apple
-
[30]
Dominique Guégan and Thomas Renault. 2021. Does investor sentiment on social media provide robust information for Bitcoin returns predictability?Finance Research Letters38 (1 2021), 101494. doi:10.1016/J.FRL.2020.101494
arXiv 2021
-
[31]
Vincent Gurgul, Stefan Lessmann, and Wolfgang Karl Härdle. 2025. Deep learning and NLP in cryptocurrency forecasting: Integrating financial, blockchain, and social media data.International Journal of Forecasting(3 2025). doi:10.1016/J.IJ FORECAST.2025.02.007
-
[32]
Suchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020. Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks.Proceedings of the Annual Meeting of the Associa- tion for Computational Linguistics(2020), 8342–8360. doi:10.18653/V1/2020.ACL- MAIN.740
-
[33]
Eric Harris-Braun, Arthur Brock, and Paul D’aoust. [n. d.]. Holochain Distributed Coordination by Scaled Consent, not Global Consensus. ([n. d.]). doi:10.1145/32 2186.322188
-
[34]
Melissa Heikkila, Chris Cook, and Clara Murray. 2025. America’s top companies keep talking about AI — but can’t explain the upsides. https://www.ft.com/con tent/e93e56df-dd9b-40c1-b77a-dba1ca01e473
2025
-
[35]
Walter Hernandez Cruz, Firas Dahi, Yebo Feng, Jiahua Xu, Aanchal Malhotra, and Paolo Tasca. 2025. AMM-based DEX on the XRP Ledger.2025 IEEE International Conference on Blockchain and Cryptocurrency (ICBC)(6 2025), 1–10. doi:10.1109/ ICBC64466.2025.11114626
arXiv 2025
-
[36]
Walter Hernandez Cruz, Kamil Tylinski, Alastair Moore, Niall Roche, Nikhil Vadgama, Horst Treiblmaier, Jiangbo Shangguan, Paolo Tasca, and Jiahua Xu
-
[37]
Walter Hernandez Cruz, Jiahua Xu, Paolo Tasca, and Carlo Campajola. 2024. No Questions Asked: Effects of Transparency on Stablecoin Liquidity During the Collapse of Silicon Valley Bank. (2024). https://arxiv.org/abs/2407.11716
Pith/arXiv arXiv 2024
-
[38]
Kia Jahanbin, Mohammad Ali Zare Chahooki, and Fereshte Rahmanian. 2023. Database of influencers’ tweets in cryptocurrency (2021-2023). 2 (2023). doi:10.1 7632/8FBDHH72GS.2
2023
-
[39]
Lily Jamali. 2025. AI firm Anthropic agrees to pay authors $1.5bn for pirating work - BBC News. https://www.bbc.co.uk/news/articles/c5y4jpg922qo
2025
-
[40]
Armand Joulin, Édouard Grave, Piotr Bojanowski, and Tomáš Mikolov. 2017. Bag of Tricks for Efficient Text Classification. 427–431 pages. https://aclanthology.o rg/E17-2068/
2017
-
[41]
Martin Juan, José Bucher, and Marco Martini. 2024. Fine-Tuned ’Small’ LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classifi- cation. (6 2024). https://arxiv.org/abs/2406.08660v2
Pith/arXiv arXiv 2024
-
[42]
Inwon Kang, Maruf Ahmed Mridul, Abraham Sanders, Yao Ma, Thilanka Mu- nasinghe, Aparna Gupta, and Oshani Seneviratne. 2024. Deciphering Crypto Twitter.Proceedings of the 16th ACM Web Science Conference, WebSci 2024(5 2024), 331–342. doi:10.1145/3614419.3644026
arXiv 2024
-
[43]
Jaehyun Kim, Thi Thu Huong Le, Sangmyeong Lee, and Howon Kim. 2024. Ethereum Smart Contracts Vulnerabilities Detection Leveraging Fine-Tuning DistilBERT.International Conference on Platform Technology and Service(2024), 133–138. doi:10.1109/PLATCON63925.2024.10830749
arXiv 2024
-
[44]
Kate Knibbs. 2025. Meta Secretly Trained Its AI on a Notorious Piracy Database, Newly Unredacted Court Docs Reveal | WIRED. https://www.wired.com/story/ new-documents-unredacted-meta-copyright-ai-lawsuit/
2025
-
[45]
Olivier Kraaijeveld and Johannes De Smedt. 2020. The predictive power of public Twitter sentiment for forecasting cryptocurrency prices.Journal of International Financial Markets, Institutions and Money65 (3 2020), 101188. doi:10.1016/J.INTF IN.2020.101188
arXiv 2020
-
[46]
Xianzhi Li, Samuel Chan, Xiaodan Zhu, Yulong Pei, Zhiqiang Ma, Xiaomo Liu, and Sameena Shah. 2023. Are ChatGPT and GPT-4 General-Purpose Solvers for Financial Text Analytics? A Study on Several Typical Tasks.EMNLP 2023 - 2023 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Industry Track(2023), 408–422. doi:10.18653/V1/2...
-
[47]
Yuan Li, Bingqiao Luo, Qian Wang, Nuo Chen, Xu Liu, and Bingsheng He. 2024. CryptoTrade: A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading.EMNLP 2024 - 2024 Conference on Empirical Methods in Natural Language Processing, Proceedings of the Conference(2024), 1094–1106. doi:10.18653/V1/20 24.EMNLP-MAIN.63 Hernandez Cruz et al
doi:10.18653/v1/20 2024
-
[48]
Zichao Li. 2025. Knowledge-Grounded Detection of Cryptocurrency Scams with Retrieval-Augmented LMs. (8 2025), 40–48. doi:10.18653/V1/2025.KNOWLLM-1.4
-
[49]
Gordon Y Liao and John Caramichael. 2022. Stablecoins: Growth Potential and Impact on Banking.International Finance Discussion Paper2022, 1334 (2022), 1–26. doi:10.17016/ifdp.2022.1334
arXiv 2022
-
[50]
Yuen C. Lo and Francesca Medda. 2020. Assets on the blockchain: An empirical study of Tokenomics.Information Economics and Policy53 (12 2020), 100881. doi:10.1016/J.INFOECOPOL.2020.100881
arXiv 2020
-
[51]
Jinghui Lu, Maeve Henchion, and Brian Mac Namee. 2020. Diverging Divergences: Examining Variants of Jensen Shannon Divergence for Corpus Comparison Tasks. (2020), 11–16
2020
-
[52]
Yichen Luo, Yebo Feng, Jiahua Xu, and Yang Liu. 2026. Resisting Manipulative Bots in Meme Coin Copy Trading: A Multi-Agent Approach with Chain-of- Thought Reasoning.Proceedings of the ACM Web Conference 2026 (WWW ’26), April 13â•fi17, 2026, Dubai, United Arab Emirates1 (1 2026). doi:10.1145/3774904. 3792635
doi:10.1145/3774904 2026
-
[53]
Sean McNally, Jason Roche, and Simon Caton. 2018. Predicting the Price of Bitcoin Using Machine Learning.International Euromicro Conference on Parallel, Distributed and Network-Based Processing(6 2018), 339–343. doi:10.1109/PDP201 8.2018.00060
arXiv 2018
-
[54]
Roberto Moncada, Enrico Ferro, Maurizio Fiaschetti, and Francesca Medda. 2024. Blockchain Tokens, Price Volatility, and Active User Base: An Empirical Analysis Based on Tokenomics.International Journal of Financial Studies 2024, Vol. 12, Page 10712, 4 (10 2024), 107. doi:10.3390/IJFS12040107
-
[55]
Satoshi Nakamoto. 2008. Bitcoin: A peer-to-Peer Electronic Cash System. 552– 557 pages. https://bitcoin.org/bitcoin.pdf
2008
-
[56]
Leonardo Nizzoli, Serena Tardelli, Marco Avvenuti, Stefano Cresci, Maurizio Tesconi, and Emilio Ferrara. 2020. Charting the Landscape of Online Cryptocur- rency Manipulation.IEEE Access8 (2020), 113230–113245. doi:10.1109/ACCESS .2020.3003370
arXiv 2020
-
[57]
Anna Paula Pawlicka Maule and Kristen Marie Johnson. 2021. Cryptocurrency Day Trading and Framing Prediction in Microblog Discourse.Proceedings of the 3rd Workshop on Economics and Natural Language Processing, ECONLP 2021 (2021), 82–92. doi:10.18653/V1/2021.ECONLP-1.11
-
[58]
Branislav Pecher, Ivan Srba, and Maria Bielikova. 2025. Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance. (11 2025), 165–184. doi:10.18653 /V1/2025.EMNLP-MAIN.9
2025
-
[59]
Pannier, Ebtesam Almazrouei, and Julien Launay
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra-Aimée Cojo- caru, Hamza Alobeidli, Alessandro Cappelli, B. Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb Dataset for Falcon LLM: Outperform- ing Curated Corpora with Web Data Only.Neural Information Processing Systems (2023)
2023
-
[60]
Arif Perdana, Alastair Robb, Vivek Balachandran, and Fiona Rohde. 2021. Dis- tributed ledger technology: Its evolutionary path and the road ahead.Information & Management58, 3 (4 2021), 103316. doi:10.1016/J.IM.2020.103316
arXiv 2021
-
[61]
Audrey Pope. 2024. NYT v. OpenAI: The Times’s About-Face - Harvard Law Review. https://harvardlawreview.org/blog/2024/04/nyt-v-openai-the-timess- about-face/
2024
-
[62]
Jacob Portes, Alex Trott, Sam Havens, Daniel King, Abhinav Venigalla, Moin Nadeem, Nikhil Sardana, Daya Khudia, and Jonathan Frankle. 2023. MosaicBERT: A Bidirectional Encoder Optimized for Fast Pretraining.Advances in Neural Information Processing Systems36 (12 2023). https://arxiv.org/abs/2312.17482v2
Pith/arXiv arXiv 2023
-
[63]
Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J
Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2019. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.Journal of machine learning research(2019)
2019
-
[64]
Mayank Raikwar, Nikita Polyanskii, and Sebastian Muller. 2024. SoK: DAG- based Consensus Protocols.2024 IEEE International Conference on Blockchain and Cryptocurrency, ICBC 2024(2024). doi:10.1109/ICBC59979.2024.10634358
arXiv 2024
-
[65]
Pornpanit Rasivisuth, Maurizio Fiaschetti, and Francesca Medda. 2024. An investigation of sentiment analysis of information disclosure during Initial Coin Offering (ICO) on the token return.International Review of Financial Analysis95 (10 2024), 103437. doi:10.1016/J.IRFA.2024.103437
arXiv 2024
-
[66]
R. L. Rivest, A. Shamir, and L. Adleman. 1978. A method for obtaining digital signatures and public-key cryptosystems.Commun. ACM21, 2 (2 1978), 120–126. doi:10.1145/359340.359342
arXiv 1978
-
[67]
Sougata Sarkar, Aditya Badwal, Amartya Roy, Koustav Rudra, and Kripabandhu Ghosh. 2025. CryptOpiQA: A new Opinion and Question Answering dataset on Cryptocurrency. 11107–11120 pages. doi:10.5281/zenodo.14469000
-
[68]
Pavlo Seroyizhko, Zhanel Zhexenova, Muhammad Zohaib Shafiq, Fabio Merizzi, Andrea Galassi, and Federico Ruggeri. 2022. A Sentiment and Emotion Annotated Dataset for Bitcoin Price Forecasting Based on Reddit Posts.FinNLP 2022 - 4th Workshop on Financial Technology and Natural Language Processing, Proceedings of the Workshop(2022), 203–210. doi:10.18653/V1/...
-
[69]
Rion Snow, Brendan O’connor, Daniel Jurafsky, and Andrew Y Ng. 2008. Cheap and Fast – But is it Good? Evaluating Non-Expert Annotations for Natural Language Tasks. 254–263 pages. https://aclanthology.org/D08-1027/
2008
-
[70]
Pollard, Eric Lehman, Alistair E
Thomas Sounack, Joshua Davis, Brigitte Durieux, Antoine Chaffin, Tom J. Pollard, Eric Lehman, Alistair E. W. Johnson, Matthew McDermott, Tristan Naumann, and Charlotta Lindvall. 2025. BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP. (6 2025). https: //arxiv.org/abs/2506.10896v1
Pith/arXiv arXiv 2025
-
[71]
Jonathan Stempel. 2025. SEC ends lawsuit against Ripple, company to pay $125 million fine | Reuters. https://www.reuters.com/legal/government/sec-ends- lawsuit-against-ripple-company-pay-125-million-fine-2025-08-08/
2025
-
[72]
Jianguo Sun, Yifan Jia, Yanbin Wang, Ye Tian, and Sheng Zhang. 2025. Ethereum fraud detection via joint transaction language model and graph representation learning.Information Fusion120 (8 2025), 103074. doi:10.1016/J.INFFUS.2025.10 3074
-
[73]
Paolo Tasca and Claudio J. Tessone. 2017. Taxonomy of Blockchain Technologies. Principles of Identification and Classification.Ledger4 (5 2017), 1–39. doi:10.5 195/LEDGER.2019.140
2017
-
[74]
Benjamin Warner, Antoine Chaffin,†Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, and Iacopo Poli. 2025. Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inferen...
-
[75]
Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Ap- pleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan Willem Boiten, Luiz Bonino da Silva Santos, Philip E
Mark D. Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Ap- pleton, Myles Axton, Arie Baak, Niklas Blomberg, Jan Willem Boiten, Luiz Bonino da Silva Santos, Philip E. Bourne, Jildau Bouwman, Anthony J. Brookes, Tim Clark, Mercè Crosas, Ingrid Dillo, Olivier Dumon, Scott Edmunds, Chris T. Evelo, Richard Finkers, Alejandra Gonzalez-Beltran,...
2016
-
[76]
Gavin Wood. 2016. Polkadot: Vision for a Heterogeneous Multi-Chain Frame- work. (2016)
2016
-
[77]
Mitchell Wortsman, Tim Dettmers, Luke Zettlemoyer, Ari Morcos, Ali Farhadi, and Ludwig Schmidt. 2023. Stable and low-precision training for large-scale vision-language models. InProceedings of the 37th International Conference on Neural Information Processing Systems (NIPS ’23). Curran Associates Inc., Red Hook, NY, USA
2023
-
[78]
Yang Xiao, Mohan Jiang, Jie Sun, Keyu Li, Jifan Lin, Yumin Zhuang, Ji Zeng, Shijie Xia, Qishuo Hua, Xuefeng Li, Xiaojie Cai, Tongyu Wang, Yue Zhang, Liming Liu, Xia Wu, Jinlong Hou, Yuan Cheng, Wenjie Li, Xiang Wang, Dequan Wang, and Pengfei Liu. 2025. LIMI: Less is More for Agency. (9 2025). https: //arxiv.org/abs/2509.17567v2
arXiv 2025
-
[79]
Yong Xie, Karan Aggarwal, and Aitzaz Ahmad. 2024. Efficient Continual Pre- training for Building Domain Specific Large Language Models.Findings of the Association for Computational Linguistics ACL 2024(2024), 10184–10201. doi:10.18653/V1/2024.FINDINGS-ACL.606
-
[80]
Jiahua Xu, Krzysztof Paruch, Simon Cousaert, and Yebo Feng. 2023. SoK: De- centralized Exchanges (DEX) with Automated Market Maker (AMM) Protocols. Comput. Surveys55, 11 (11 2023), 1–50. doi:10.1145/3570639
doi:10.1145/3570639 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.