Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:16:20.792657Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2501.10322.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:16:20.792657Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T03:42:21.307552Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:47:48.559145Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3de18d85-dfa7-4811-aa6a-c90d5856cd2a · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbeb6181-a3c8-489b-9edd-16cc8ea47f3f · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har2 S ubword: Extending the subword embedding space using robust character compositionality
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 933d638a-7d56-4add-b6e5-84d415df53d7 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Character-level language modeling with deeper self-attention
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb07750b-a32c-4997-a03a-f78160021689 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ece48a5-8ffd-460b-bafa-2bb2cc5d30a1 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Semantic parsing on F reebase from question-answer pairs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1220c9c1-7765-4959-b42b-3d103c8aa7e6 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Piqa: Reasoning about physical commonsense in natural language
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 414fed55-78ec-4436-92e8-2150c3d30aa3 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Occiglot fineweb v0.5, 2024
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0102fa2-af1d-4228-8664-4be39b1f8aa5 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bridging the Gap for Tokenizer-Free Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe2a8b4-c99a-43ee-a251-0ea9812ca2f9 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B ool Q : Exploring the surprising difficulty of natural yes/no questions
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e54a72f7-d208-4f53-ba61-ca3ff6cbd064 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Bowman, Holger Schwenk, and Veselin Stoyanov
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9904d11-5bd0-4444-b023-d8eae768ce5a · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher R \' e
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a168cc1-d65a-4b9d-aa61-d38ce01d6f32 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T-FREE: Subword Tokenizer-Free Generative LLMs via Sparse Representations for Memory-Efficient Embeddings
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 782f88d9-861e-4520-b119-60c30e302157 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models BERT: pre-training of deep bidirectional transformers for language understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da89344-4a69-4cd2-823b-bccc91664fb0 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b555442-5b81-4784-aa2a-f7b709d46e53 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C haracter BERT : Reconciling ELM o and BERT for word-level open-vocabulary representations from characters
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73b683fb-60de-4cce-92ee-48bb9f33c1c8 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A new algorithm for data compression
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 656f9556-601f-446c-b152-fdca2744171e · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models A framework for few-shot language model evaluation, 07 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b78784f5-7632-4cdf-ad9a-b779d1a3f5eb · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Better & faster large language models via multi-token prediction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9eca35e-c863-4917-81db-3df6e9cd72de · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Botvinick, Ian Simon, Hannah Sheahan, Neil Zeghidour, Jean - Baptiste Alayrac, Jo \ a o Carreira, and Jesse H
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f10e283-7a25-418f-b2b8-b3659d8ce36a · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Pile of Law: Learning Responsible Data Filtering from the Law and a 256GB Open-Source Legal Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93bdbdc7-51e1-44f8-a080-af2eab71825a · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Measuring massive multitask language understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88d0a68e-a3a3-4cf2-8383-2381420f2759 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models T rivia QA : A large scale distantly supervised challenge dataset for reading comprehension
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be429736-ed51-4512-9f19-57a2ea1d3b50 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models DataComp-LM: In search of the next generation of training sets for language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 660921d5-e930-4439-9f8b-10154c24dde9 · outbound
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8933bb48-3fbe-4f11-8f7c-4ef8681aafba · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Truthfulqa: Measuring how models mimic human falsehoods
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ecbde2d-a12a-4623-81e2-61804aaee209 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Decoupled weight decay regularization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa6d751-0c1e-498a-89de-b04d24e64ae2 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models C har BERT : Character-aware pre-trained language model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96634558-8a12-4a1e-883d-2375d3d142b5 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbac2ce3-d16b-4cb6-a7e4-e25169452ad4 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7114735-6bee-4cc9-a853-5bd59a842d07 · outbound
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fb4d1ed9-bada-4e93-8008-4a8e2c546bca · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4920c0d1-870a-4343-9e07-a1ec82e2e7b6 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Openwebmath: An open dataset of high-quality mathematical web text, 2023
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f524630-137f-4cae-ad5c-b0dfdde3b122 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b122c748-f7d3-4765-b9f1-722e781c350d · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Language model tokenizers introduce unfairness between languages
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0087ec6c-2aa5-4ee4-9ba8-bd65f7aed6bc · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models W i C : the word-in-context dataset for evaluating context-sensitive meaning representations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336de9ed-8e9f-4f7a-8c1e-2ee9afbacf21 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Germanbenchmark, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 38759c8e-ccbc-435a-aa92-a26b40a9f5cc · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Efficiently scaling transformer inference
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 169d0db1-22b1-4897-958f-6eeaf8fe2c05 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Winogrande: an adversarial winograd schema challenge at scale
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5f9226-e877-49ac-b01d-6c3740db7e4c · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Neural machine translation of rare words with subword units
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20926a3b-0f62-45e6-a6f7-e97a45f61b93 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models SpaceByte: Towards Deleting Tokenization from Large Language Modeling
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4a657ca-7ccf-4c2b-b7eb-eefefbfce51c · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models From characters to words: Hierarchical pre-trained language model for open-vocabulary language understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3afa7e17-d33e-4208-9623-f0e9d23d5e58 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Tran, Sebastian Ruder, Jai Prakash Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e6ffb928-35e1-4086-96e2-56637c7195a6 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Learn your tokens: Word-pooled tokenization for language modeling
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878598da-bec5-4c1e-8e7e-89b674f45275 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7d60f3f-f0d0-4a13-9bd3-73083189e0f4 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Skywork: A more open bilingual foundation model, 2023
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 937b5811-b6f4-445d-ba3f-d3069313addb · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models B y T 5: Towards a token-free future with pre-trained byte-to-byte models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ac978a-6974-4c1c-aac0-18890fd35edd · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models MEGABYTE: predicting million-byte sequences with multiscale transformers
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a9035346-a518-4497-9d04-4106f6631988 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8770478-1b87-44cb-9c73-05f7bd567182 · outbound
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f090c69b-7d44-49b4-9c00-dfd4e5a5f5d7 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce16206-2216-4b9e-9bdd-b8a2ac8537d4 · outbound
Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b73c34-d868-473f-b1c9-35637e98a049 · inbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8705c67-22f6-4b94-8b48-f2e04376b670 · inbound
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.