Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:05:28.713898Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2605.09630.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:05:28.713898Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 104 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a5b37670-bc1b-42e1-9b35-d2fcf54fbf07 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df4ae625-4042-4327-b252-9c5c2f7e1cd0 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Character-level language modeling with deeper self-attention
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d39d824f-347c-4ed5-8e2d-21940bc8b05e · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Program Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bed9bb2b-1293-40a8-904e-24184a3745c3 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Relaxed recursive transformers: Effective parameter sharing with layer-wise lo RA
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ccde85b-60f6-435b-b7af-0840607b71f9 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eafd033d-358c-4001-8d82-9f8e21d16ff0 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Pondernet: Learning to ponder
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4282b008-a885-4b7c-ab7e-b829f150f779 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Large Concept Models: Language Modeling in a Sentence Representation Space
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17154d6a-e643-4fcb-8200-4c6b80ffbb30 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael K Kopp, G \
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 949ad529-ec5c-48ed-a348-2fbcaac31232 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Conditional Computation in Neural Networks for faster models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0bab88ea-4fe3-4761-9724-65e1cded232b · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models A neural probabilistic language model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fdae45e8-ae3d-465c-b256-6f2d42879c7d · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Piqa: Reasoning about physical commonsense in natural language
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6a01556e-eff6-4dfc-8640-16dcbe040bbe · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Language models are few-shot learners
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae5aef9f-eea5-4b88-acf7-747a27194c85 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lee, Deming Chen, and Tri Dao
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2ef76be8-a813-47ca-819b-d8b1dd0aa5bf · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Evaluating Large Language Models Trained on Code
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6e6d25bf-49f9-4dfa-9e76-5b7fb45eac1b · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Bridging the Gap for Tokenizer-Free Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e9e5463b-26bd-4b2f-872b-01021a4a18c8 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical multiscale recurrent neural networks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d9b151f9-1c12-4ae0-9a15-1b3eafd9a618 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models B ool Q : Exploring the surprising difficulty of natural yes/no questions
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b01e1819-f1a9-48f4-9c02-ada5a17b18e3 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Clark, Dan Garrette, Iulia Turc, and John Wieting
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3b7a85e4-e7d7-4162-87c2-e766bcdbc3c0 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb5807a7-a5f9-4d7a-a453-3ad7ae5517fa · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models No Language Left Behind: Scaling Human-Centered Machine Translation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ade6638-ef35-4795-83e4-6b3a6dcee219 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mo EUT : Mixture-of-experts universal transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fa4007d-2fa5-402c-a808-8700e9a1430f · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Getting the most out of your tokenizer for pre-training and domain adaptation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation db0c604f-3d80-454a-89f7-4c87eb1f8c52 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Funnel-transformer: Filtering out sequential redundancy for efficient language processing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 104d1fa3-97d4-4013-825f-852973f78e50 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9484715a-1b7d-40bb-b06d-e411caf3904a · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Universal transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e0e1c019-e0d4-4742-88b2-d410cf5cdd24 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b19ad1c0-9ce3-4c20-a8d0-282dd964c99d · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models A new algorithm for data compression
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9138ff7f-3b7f-41bf-8e83-4b0a7ea88ed9 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Improving Language Understanding from Screenshots
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3529846b-e939-45e3-a587-7da592bd7876 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8741a2a4-54c7-4e18-a16b-c2529ea277f1 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lee, and Dimitris Papailiopoulos
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3ebf462-cd4d-47b4-ad71-f1d60bfd6ba2 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Better & Faster Large Language Models via Multi-token Prediction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 533c8037-8125-47ae-bd69-a95e6ce28393 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MANT a: Efficient gradient-based tokenization for end-to-end robust language modeling
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a10a360-d88d-46e3-b34f-e83cc777fd41 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d1635a5-75c3-4166-8d12-393f709b1fec · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Think before you speak: Training language models with pause tokens
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 448c6cfa-953e-4975-a68c-a7094cba5aa4 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Generating Sequences With Recurrent Neural Networks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a3b8d8a6-1cfa-49c3-8122-20fda793484b · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Adaptive Computation Time for Recurrent Neural Networks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c93042f3-5792-45b5-bacd-e505f59884de · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Fast and expressive multi-token prediction with probabilistic circuits
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df9c3301-d60b-4170-9348-944cce3b88c7 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mamba: Linear-time sequence modeling with selective state spaces
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e179736f-a2fd-408d-afde-1ea02f97e1af · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models OLMES: A Standard for Language Model Evaluations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 52d0423a-5349-463c-b03d-a209ee79f5e3 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bd8a292b-ab5b-435c-970d-6f08198a24d0 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models General-purpose, long-context autoregressive modeling with perceiver AR
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0d9e771a-d6fa-4a48-88ec-1e9de7d0b470 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Measuring massive multitask language understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6fb07288-5514-4f0a-ab1c-cae81b06a61c · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Block transformer: Global-to-local language modeling for fast inference
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 763c9027-4771-4aaa-9e7a-596e99e52afd · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Deep networks with stochastic depth
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29771bf6-6e85-4dd7-b0a8-9a157412cf50 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Conceptmoe: Adaptive token-to-concept compression for implicit compute allocation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c19a60f8-b258-42c8-85ac-9e512f693668 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Character-level language modeling with hierarchical recurrent neural networks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c9d4a8af-d64c-413b-9b18-497755608a41 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2fa510aa-69a7-4e8d-bc35-d13d9db76557 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 32ec8fda-688a-48da-a1da-2688492f6782 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Perceiver: General perception with iterative attention
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 303385c4-0d43-4ac6-b8f8-3d9d6249d87c · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models `` low-resource '' text classification: A parameter-free classification method with compressors
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b180b5a6-dc70-4904-b5e7-e5b1062322d6 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MrT5: Dynamic Token Merging for Efficient Byte-level Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation abffccac-c31b-4d9e-902c-40b6ea87487d · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Subword regularization: Improving neural network translation models with multiple subword candidates
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3d2df5aa-93d6-40e6-b645-f69dba7fb476 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9e563b4f-ee86-416c-bfb6-345a22a3786a · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mamba-3: Improved sequence modeling using state space principles
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7861d0da-8c14-4247-a758-834f7c2ff845 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2f26fe72-8ceb-49ad-9377-9faedc558799 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Training LLMs over Neurally Compressed Text
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1d5eba93-1f43-4c78-9e57-d6d325bc3670 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models DataComp-LM: In search of the next generation of training sets for language models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9699f24d-0927-420a-a535-75180591fe3d · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b91d14af-756d-4463-bb37-1c242002f7aa · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Smith, and Yejin Choi
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 29744b20-b3d1-42b1-ae19-f3cb4d3bef1f · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Decoupled weight decay regularization
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b20de449-b3d5-4d72-8919-2c90cc6c6d94 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Text rendering strategies for pixel language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d928f038-711e-4ed1-b0cd-7249749f9f33 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Starcoder 2 and the stack v2: The next generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4ecb3e0d-da08-4a44-86df-edf53afc970e · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models The art of prompt design: Prompt boundaries and token healing
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 319fe42b-e335-4e72-b036-6550ee642834 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Guidance
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 323195a1-9d0d-4ac9-9322-5e67180d5615 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b6b19845-325a-49f3-846d-2831fa844e7b · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Minixhofer, T
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f2e5a378-ef06-4093-8ff7-0a407bcee020 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical transformers are more efficient language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c428cfe-862d-4929-a668-421fed0f9a22 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient transformers with dynamic token pooling
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2b73c34-d868-473f-b1c9-35637e98a049 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2a0d88a2-9ad4-4115-af90-bdca6c3c9326 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models GPT-4 Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cd7fb21d-0d92-40c9-b092-663f8dacf775 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models FLEXITOKENS: Flexible Tokenization for Evolving Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 53a577bf-b816-4a10-ad63-442b74c0309b · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Byte Latent Transformer: Patches Scale Better Than Tokens
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 626da5f5-2f83-4268-93c0-77ac5a6a2eaa · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Openwebmath: An open dataset of high-quality mathematical web text
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5d331cc4-1b11-4e5f-be0f-46429cf8d53c · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Dynamic large concept models: Latent reasoning in an adaptive semantic space
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ea47df25-b79d-4ca1-9240-1bf4cbffcfb2 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Learning to Generate Reviews and Discovering Sentiment
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 73f206ef-6d41-4de0-adb0-ac5abc801272 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04407f9a-01d5-44cf-9868-0dd310ee8d78 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a776c116-46c5-412f-a994-dc32c9c80966 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Solidgoldmagikarp (plus, prompt generation)
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7f05fa1-038f-42ef-82d5-bf77bda3a08f · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 75c4f69a-aa95-4c86-be27-ff4bfa0b4f73 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Winogrande: An adversarial winograd schema challenge at scale
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9c91c57b-26cb-46b0-b7b0-7aea5313010c · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Robust open-vocabulary translation from visual text representations
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1358df84-fcff-433a-8038-2887c466d418 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Japanese and korean voice search
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0cae3c0e-7e90-4dea-9d85-9c8524f753b2 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Neural machine translation of rare words with subword units
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b5ee4b2-cd9a-4ea0-b7ab-431894f91dc0 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models SpaceByte: Towards Deleting Tokenization from Large Language Modeling
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cb630b05-c1e7-4b16-8bd7-fa16816d1380 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Blockwise parallel decoding for deep autoregressive models
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4662c9dc-dd95-470e-bfde-e5aed21e6f98 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Generating text with recurrent neural networks
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 818ac257-47a3-4845-b669-46133384fa50 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Sparse universal transformer
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3fd762a8-b3b8-45ab-9edf-0c209165ee4d · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b906c358-94d1-46ea-9b78-cf7d5c7a9722 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models From Bytes to Ideas: Language Modeling with Autoregressive U-Nets
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation eb950473-1012-45be-9eac-ab678a94bd2e · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MambaByte: Token-free Selective State Space Model
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1aed69f3-a39b-4bf3-9295-2f7d1846d93f · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Parallel loop transformer for efficient test-time computation scaling.CoRR, abs/2510.24824
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f683faf2-3189-4eec-9029-366a612255d5 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient Pretraining Length Scaling
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c881321-048e-4011-9d33-d52fbfe35ff7 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc9d9a3b-a53c-445e-a988-fe3fd882444f · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models B y T 5: Towards a token-free future with pre-trained byte-to-byte models
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation daccf9ec-9873-4498-94a3-b761c17004b8 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Problematic Tokens: Tokenizer Bias in Large Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 77da302a-c137-4766-8f78-592c9548acb2 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Scaling embedding layers in language models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 17cbaf07-d478-487a-9ddb-132efd34862c · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MEGABYTE : Predicting million-byte sequences with multiscale transformers
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ddd6ee95-0fdc-4e84-9b82-880d6749e8ac · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7d255cad-4e4f-4144-8e01-2a0a326d65e1 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Ponder LM : Pretraining language models to ponder in continuous space
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f24ce0fc-075b-4fac-b0fc-10ba1da542c1 · outbound
Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Linear complexity randomized self-attention mechanism
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
No inbound Pith citation observations are available.