Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:39.195688Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 44 inbound Pith citation observations for arXiv:2412.09871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:39.195688Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:33:18.628297Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:47:48.461365Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1d369c27-aa05-4d61-af2b-cfa6d1fbc0a6 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 693a9cad-1457-40e7-a424-220f9a8df050 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Character-level language modeling with deeper self-attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cba0e444-0c7b-49fd-9f14-e15af1736dcf · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Program synthesis with large language models, 2021
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c30d31-d107-4ede-8a6b-85e9ee2efc52 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Learning to rank with (a lot of) word features
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cf52b182-05b7-4f9d-ad8b-e0ead61a17f0 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Piqa: Reasoning about physical commonsense in natural language
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a8fe8fa5-940d-4194-9c0c-affa3b7d8ae8 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Transformer flops, 2023
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5b7afab3-5a82-4bb1-81d1-6e72a14c32a8 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dc9a6f5-578e-466b-ac02-96b8bbde30ce · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Bridging the Gap for Tokenizer-Free Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce895d6c-f361-463f-bf89-011fb947120d · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Hierarchical multiscale recurrent neural networks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4b746304-68f1-4a60-ae44-b8c384b897fe · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Canine: Pre-training an efficient tokenization-free encoder for language representation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0cdafcf7-2570-430d-b90b-f30140e021a5 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Think you have solved question answering? T ry ARC , the AI2 reasoning challenge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 14bcd589-c480-4199-b336-ab787b060119 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Getting the most out of your tokenizer for pre-training and domain adaptation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b386dd3b-498c-4a71-b9c8-81d30089216d · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Flash A ttention: Fast and memory-efficient exact attention with io-awareness
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fe1d1f3-2faa-48e1-a66e-9903e32e0482 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens The llama 3 herd of models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f49ea3-17dc-469a-a24c-c31278ade781 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens CUTE : Measuring llms' understanding of their tokens
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b5acf611-f4dc-492a-94e7-45bb832d4937 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens CharacterBERT : Reconciling elmo and bert for word-level open-vocabulary representations from characters
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 87143f23-eccb-4df1-88fb-bde3c2bab50c · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens A new algorithm for data compression
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d71886da-cd17-4312-820e-719cd3150153 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens The F lores-101 evaluation benchmark for low-resource and multilingual machine translation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db7ffea0-0c71-4771-85c2-a2db1606698b · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Generating sequences with recurrent neural networks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5aa62707-6ff5-45ed-9ee7-81a236097e48 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Mamba: Linear-time sequence modeling with selective state spaces
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 72238b77-86f6-45d8-9856-b16fe44bfff0 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Measuring massive multitask language understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 380e68d6-1f27-4146-87f7-dfdc9a458a85 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Training compute-optimal large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5f0a86ce-4196-445b-8f88-ee07e797f445 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Perceiver: General perception with iterative attention
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3c809974-f515-4f4d-8744-5d2a407f73f5 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Neural machine translation in linear time
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 967658ef-d4ac-4249-af3a-a6d12550dd46 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Scaling laws for neural language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e24b8f0-ff74-4360-8cf5-7b865c2555ea · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Byte-level machine reading across morphologically varied languages
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b0f522bc-7e01-4e2f-b88b-5a0b38957714 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Character-aware neural language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 730d1e14-8232-447a-bc9a-711598e7c305 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Training llms over neurally compressed text
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d2ad2aee-659e-4ee8-939f-55047a6d918b · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Datacomp-lm: In search of the next generation of training sets for language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f9ff057d-ea1f-4852-b6a7-90c101f4d4dc · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Xlm-v: Overcoming the vocabulary bottleneck in multilingual masked language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2293348e-1ad8-4bbd-8279-5515acc108ed · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Myte: Morphology-driven byte encoding for better and fairer multilingual language modeling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f0076ab-9dc7-4f70-b862-4e8ee5c97ee2 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Decoupled weight decay regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa05016-eeb1-4dca-b62f-f56c836db36c · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Subword language modeling with neural networks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fbb644b8-3b75-4b2a-ab6f-c2cc47dc369a · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Hierarchical transformers are more efficient language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4fe4ac8f-66a5-4b9c-93c9-a9984becc668 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Efficient transformers with dynamic token pooling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 485d1b9b-f9a8-4122-bcfd-8a7eb38910f2 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Language model tokenizers introduce unfairness between languages
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ca3519ec-3e35-427d-934c-578e910b3fb3 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Language models are unsupervised multitask learners
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbcbaada-7741-4cfa-8d9b-b774b8062ec3 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Neural machine translation of rare words with subword units
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 510b83f9-51a5-4cda-8cda-0c8543e7a90a · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens GLU variants improve transformer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e1a77165-4065-4ff2-baaf-771e5db04329 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Spacebyte: Towards deleting tokenization from large language modeling
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 17801808-54a4-4813-85f8-185c26876ab7 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens RoFormer : Enhanced transformer with rotary position embedding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation caf19b92-2190-4463-b62f-5ade5d799632 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Generating text with recurrent neural networks
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 270c9470-2275-4a8f-b6d5-999758b15a60 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Phonologybench: Evaluating phonological skills of large language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3ca97346-4abe-4d7b-8636-46d0eecc1035 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Llama 2: Open foundation and fine-tuned chat models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5a9268c6-dc49-4989-a1ab-829d708d70fc · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d045f22b-b877-4600-8092-1d1e512f0f5f · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Mambabyte: Token-free selective state space model
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0a1bb76-62f3-451a-ab6a-9d47929cc9a8 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Effective long-context scaling of foundation models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5a8c4775-7903-4985-a682-18a2366f5f28 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Byt5: Towards a token-free future with pre-trained byte-to-byte models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48951ae3-14b0-4eb1-b4c8-d616b85b24eb · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Megabyte: Predicting million-byte sequences with multiscale transformers
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a695aa24-681d-451a-8939-add0cd7362bf · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Hellaswag: Can a machine really finish your sentence? arXiv, 2019
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fdbcdc63-6a07-4fc5-8fe5-83cbfe55d578 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Root mean square layer normalization
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7d0a80c5-90c5-4e1b-9c44-a43cfab602d5 · outbound
Byte Latent Transformer: Patches Scale Better Than Tokens Character-level convolutional networks for text classification
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8a56eef1-ae4b-46b4-907c-c68d04da976d · inbound
Tokenisation is NP-Complete Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4326a6ca-de80-49ce-a5ac-fd0fff6d75e1 · inbound
Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126eaf8b-4455-492c-9284-6290608b28d0 · inbound
Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d88bde0-ea82-48b9-b831-7338928bc81f · inbound
LLM-based event log analysis techniques: A survey Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b92a72-72c2-4424-bec4-17a25157877c · inbound
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb9f469-f237-4676-8224-e6cc4e59f5f3 · inbound
Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f1b5c5-f89a-4c3c-a615-48294aef94ca · inbound
An Uncertainty Principle for Linear Recurrent Neural Networks Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca9a8266-a026-4ed1-be69-2768a92a6298 · inbound
ALFEE: Adaptive Large Foundation Model for EEG Representation Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a162554-a44c-43b1-b0a0-d211fd6668d2 · inbound
FreeMesh: Boosting Mesh Generation with Coordinates Merging Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f9c47a7-419b-464e-a43d-3d6ef8f0130a · inbound
EXECUTE: A Multilingual Benchmark for LLM Token Understanding Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48674664-e288-4d0f-a3a2-a70e1a776f46 · inbound
Improving Language and Modality Transfer in Translation by Character-level Modeling Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1085e6-f3aa-4d90-8b7f-dc8d684d1442 · inbound
E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435b1c62-d240-45d8-be4e-5163e52aa492 · inbound
Sampling from Your Language Model One Byte at a Time Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13953cd7-5da7-420e-bdb4-4ca577edf254 · inbound
From Bytes to Ideas: Language Modeling with Autoregressive U-Nets Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ade3c6-115f-48e0-99a6-2e3a237063e1 · inbound
Entropy-Driven Pre-Tokenization for Byte-Pair Encoding Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 1999
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1db18ae7-1465-44d1-bff7-ba515ca024bd · inbound
ByteSpan: Information-Driven Subword Tokenisation Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6722b8-8bcf-42d5-ba56-645dc9e209fc · inbound
Learning to Skip the Middle Layers of Transformers Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baca7f72-cda5-452d-9c71-b893ffbde6cf · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 280
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62add196-6752-4877-bae9-35fb7daa884e · inbound
Energy-Based Transformers are Scalable Learners and Thinkers Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e67bcc0-2611-495d-9d96-e7291394e431 · inbound
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13ab9f17-58fd-4f22-91c6-1c85f375d120 · inbound
FLEXITOKENS: Flexible Tokenization for Evolving Language Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b6c96cd5-26b4-46c5-a8f7-55e1c4d53548 · inbound
Synergy: End-to-end Concept Model Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245cef04-c4a6-4041-b171-e7fca218e328 · inbound
SpeLLM: Character-Level Multi-Head Decoding Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38465e18-0e65-420d-a0d2-dcc00bdde284 · inbound
Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a868451e-b1a5-421b-aa4e-550ebee901fb · inbound
Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 800d5deb-5edf-4a87-b26e-47a6a5381726 · inbound
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 91b7b0a1-0333-4256-9c9d-5b51041eeb5e · inbound
Lost in State Space: Probing Frozen Mamba Representations Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 53a577bf-b816-4a10-ad63-442b74c0309b · inbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7df2ea19-0fd7-49aa-9205-13bf95949310 · inbound
Bin Latent Transformer (BiLT): A shift-invariant autoencoder for calibration-free spectral unmixing of turbid media Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b84543b0-8be2-479c-9859-cf5a6484eceb · inbound
Towards Understanding Self-Pretraining for Sequence Classification Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba6156fc-1975-4c98-86fe-078c7e3ebb5c · inbound
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3c62183c-42a5-4164-9825-5951dd2d4e5b · inbound
Large Byte Model: Teaching Language Models About Compiled Code Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation beafb45c-a498-40d1-a858-65063d412e83 · inbound
MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9283bf06-bb4d-4426-a496-a5bb87b54656 · inbound
Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 22acbaf7-26c8-4e4d-93bd-f5e7789f4c1e · inbound
User as Engram: Internalizing Per-User Memory as Local Parametric Edits Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b81a4b1-4169-4196-acf0-2eccc0eae93e · inbound
Protocol-Aware Tokenization and Architecture Co-Design for Wireless Packet Foundation Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9cc5f5e6-a47f-4b00-8cf0-083dd64b678e · inbound
Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2cdd0554-5906-4ba3-933a-f63d452445c0 · inbound
EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 13391aad-58d9-4548-a284-6bc4e487ded6 · inbound
Cybersecurity is the True Frontier for Generative AI Success or Failure Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c451fb12-4662-40ad-9d8e-929f82ac05e3 · inbound
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8755bc1c-bb44-4e9b-a4f2-7bf298987e22 · inbound
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d66f14-9838-4e25-9a40-c438a24c5c58 · inbound
Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0799d3ef-482e-4656-ad44-e5616d0dca69 · inbound
Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d6c47de-2dd6-444f-a2bb-d543f28ab848 · inbound
Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.