Pith. sign in

Paper Citation Record · LEDGER

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

As of 6 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 0 inbound Pith citation observations for arXiv:2605.09630.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09630 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:05:28.713898Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 104 outbound references displayed

  • verified exact31
  • verified fuzzy58
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch10

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a5b37670-bc1b-42e1-9b35-d2fcf54fbf07 · outbound

This paper cites MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MAGNET: Improving the Multilingual Fairness of Language Models with Adaptive Gradient-Based Tokenization

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.152842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:0812ed21500f0db91055b2b125efaf8a3e597935fed42a37676bd24c5f508ab5

Observation df4ae625-4042-4327-b252-9c5c2f7e1cd0 · outbound

This paper cites Character-level language modeling with deeper self-attention.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Character-level language modeling with deeper self-attention

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.477886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:29c23e1832355a15d6041376440f9056ca9bf7cd1b0feeada6e4250600f0e3a9

Observation d39d824f-347c-4ed5-8e2d-21940bc8b05e · outbound

This paper cites Program Synthesis with Large Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:29.307355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:a9d967d630f80921a5a6ab8d306d47a177eb39d8da273288b5f636ecc77b58ac

Observation bed9bb2b-1293-40a8-904e-24184a3745c3 · outbound

This paper cites Relaxed recursive transformers: Effective parameter sharing with layer-wise lo RA.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Relaxed recursive transformers: Effective parameter sharing with layer-wise lo RA

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.446606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:6bf9d3ff6952393f13c666c8a547d637aeaa4ccefeac699cee621f000b34e39b

Observation 4ccde85b-60f6-435b-b7af-0840607b71f9 · outbound

This paper cites Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mixture-of-recursions: Learning dynamic recursive depths for adaptive token-level computation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.489276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:64ea28541b7ea543fab5c31bc37e23296c001de018291cd38433853da1377187

Observation eafd033d-358c-4001-8d82-9f8e21d16ff0 · outbound

This paper cites Pondernet: Learning to ponder.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Pondernet: Learning to ponder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.318848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:eb0119a79c52e8a0a4d635bb7886309efbdab89148a0d74eabac003579804b57

Observation 4282b008-a885-4b7c-ab7e-b829f150f779 · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.938338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:90ec010d82d90e2c95971e986bcfcab1f5026459f86ee97b10e9cefa2dcb7269

Observation 17154d6a-e643-4fcb-8200-4c6b80ffbb30 · outbound

This paper cites o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael K Kopp, G \.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael K Kopp, G \

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.469235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:9f9afe24fbb4c835f059eb1605993937e5656b1e64dc5e2af325ffac2df03b68

Observation 949ad529-ec5c-48ed-a348-2fbcaac31232 · outbound

This paper cites Conditional Computation in Neural Networks for faster models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Conditional Computation in Neural Networks for faster models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:41:24.655364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:62c1e0525a1cb1e662ef8f56e1ebf7a40df6036606dac7bcc1523bfba4082dbb

Observation 0bab88ea-4fe3-4761-9724-65e1cded232b · outbound

This paper cites A neural probabilistic language model.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models A neural probabilistic language model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.293983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:7d4fa76a3c8b566ba05862d3e0e7c35ebca4f9e11c6334bbef83d08679e20dc2

Observation fdae45e8-ae3d-465c-b256-6f2d42879c7d · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Piqa: Reasoning about physical commonsense in natural language

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.307227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:620f8024dcc2030e301a90ee1e6d1c489df77b05e6882a6e6b782f5674989ea7

Observation 6a01556e-eff6-4dfc-8640-16dcbe040bbe · outbound

This paper cites Language models are few-shot learners.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Language models are few-shot learners

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.389613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:c488ed6849ac4448c8499393e5e6bde21eb688443be4c4c8ec804a2dfe7db985

Observation ae5aef9f-eea5-4b88-acf7-747a27194c85 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lee, Deming Chen, and Tri Dao

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.396132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:48fa67051a89042908c9b1107fb6654e3fdd6c7cad9ed5774b3a9b996768c7f8

Observation 2ef76be8-a813-47ca-819b-d8b1dd0aa5bf · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Evaluating Large Language Models Trained on Code

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:28.849357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:01be3def7acef6913b23509c147cee0653d7b9cd42fdbc494a5354ce07d19bab

Observation 6e6d25bf-49f9-4dfa-9e76-5b7fb45eac1b · outbound

This paper cites Bridging the Gap for Tokenizer-Free Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Bridging the Gap for Tokenizer-Free Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.262266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:a92c3a82ffd61c0067c48a9fbdaf247eb480ec412ec1e769ecae7b21ac2ffedb

Observation e9e5463b-26bd-4b2f-872b-01021a4a18c8 · outbound

This paper cites Hierarchical multiscale recurrent neural networks.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical multiscale recurrent neural networks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.310954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:e0f69011d16835f5be15b355ef83b05e5c9201fe206f1b70afbf0d7aa1b6db77

Observation d9b151f9-1c12-4ae0-9a15-1b3eafd9a618 · outbound

This paper cites B ool Q : Exploring the surprising difficulty of natural yes/no questions.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models B ool Q : Exploring the surprising difficulty of natural yes/no questions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.392949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:0cb94fc1a3dbe6ad92d23dbc0c17c08664eb0ce7d137a7fd76a2e5df6a7ac241

Observation b01e1819-f1a9-48f4-9c02-ada5a17b18e3 · outbound

This paper cites Clark, Dan Garrette, Iulia Turc, and John Wieting.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Clark, Dan Garrette, Iulia Turc, and John Wieting

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.507672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4a6167dea8b44cb10306e8144fb58edf2cc615bd0c64ff1cba281c9935873502

Observation 3b7a85e4-e7d7-4162-87c2-e766bcdbc3c0 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:29.139734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4a168a8056af71b0a7872f71233566ff01063d05662c2aef12f3fbb4be119d46

Observation eb5807a7-a5f9-4d7a-a453-3ad7ae5517fa · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:54:14.284800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:17ec89ecdea4da7823cd14360189cc50121ca219d89d82f87a7f85139b52ac49

Observation 7ade6638-ef35-4795-83e4-6b3a6dcee219 · outbound

This paper cites Mo EUT : Mixture-of-experts universal transformers.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mo EUT : Mixture-of-experts universal transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.458156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:1e8bb78afd18727ec1b37d87605c9259f0367de2f1008cb5f36b65b23290a5d1

Observation 3fa4007d-2fa5-402c-a808-8700e9a1430f · outbound

This paper cites Getting the most out of your tokenizer for pre-training and domain adaptation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Getting the most out of your tokenizer for pre-training and domain adaptation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.790238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:49980dcf5ba941a62f6d8a95e3966c91cf6d75ca5b267422e6332a421a3d4592

Observation db0c604f-3d80-454a-89f7-4c87eb1f8c52 · outbound

This paper cites Funnel-transformer: Filtering out sequential redundancy for efficient language processing.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Funnel-transformer: Filtering out sequential redundancy for efficient language processing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.298623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:8b6f1e646911b9b2b4c4e965f0adafe90ed767c485ae55defbf466ea7deb2dcc

Observation 104d1fa3-97d4-4013-825f-852973f78e50 · outbound

This paper cites Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Transformers are SSM s: Generalized models and efficient algorithms through structured state space duality

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.442994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:383688f7e8dfd72dd49d266bd5586dbe232036eaac5525cb24eb5242ef9b9891

Observation 9484715a-1b7d-40bb-b06d-e411caf3904a · outbound

This paper cites Universal transformers.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Universal transformers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.402438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:41fd46e3a59b03a52996300d04244b208bbac98b654981418a187083e222b4bf

Observation e0e1c019-e0d4-4742-88b2-d410cf5cdd24 · outbound

This paper cites BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 26

Resolution
verified exact
doi, observed 2026-05-12T04:06:19.682237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:12dba9c490e615feef069ff338a11951367fc159848410b5ef77e0340b103e63

Observation b19ad1c0-9ce3-4c20-a8d0-282dd964c99d · outbound

This paper cites A new algorithm for data compression.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models A new algorithm for data compression

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.379818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:7c56ff345a9e2f6ef942e3005754f9fc483d269ef0f10e4017fbd88bc0fa6b57

Observation 9138ff7f-3b7f-41bf-8e83-4b0a7ea88ed9 · outbound

This paper cites Improving Language Understanding from Screenshots.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Improving Language Understanding from Screenshots

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.128760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:a473c31db6ec03df5e5de15e95fc078a49dd9a4823634c0175d1a5c50219188b

Observation 3529846b-e939-45e3-a587-7da592bd7876 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:39:41.804715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:8d631a4c5b28a935a8e595370af62b08e9cdca1cc5a599192915365af38c8095

Observation 8741a2a4-54c7-4e18-a16b-c2529ea277f1 · outbound

This paper cites Lee, and Dimitris Papailiopoulos.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lee, and Dimitris Papailiopoulos

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.484868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:e5486109b303942e0ce3e49b9f5157f3d09855400029627f6f6bb219c5da86e8

Observation a3ebf462-cd4d-47b4-ad71-f1d60bfd6ba2 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Better & Faster Large Language Models via Multi-token Prediction

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:26:09.884823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:6a852ec992bfad2420fd05e3bd1dbab7f3511d7679944d15b9d2e365b726ab61

Observation 533c8037-8125-47ae-bd69-a95e6ce28393 · outbound

This paper cites MANT a: Efficient gradient-based tokenization for end-to-end robust language modeling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MANT a: Efficient gradient-based tokenization for end-to-end robust language modeling

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.461769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:492649973220b36e3ac8f69f546da7036d05a3ab4765ec8fc894837cc81f529a

Observation 9a10a360-d88d-46e3-b34f-e83cc777fd41 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:29.003298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:21caae11d38dc5d7358c00f7b01418339e32acffce7bac6848667f602f5fba87

Observation 1d1635a5-75c3-4166-8d12-393f709b1fec · outbound

This paper cites Think before you speak: Training language models with pause tokens.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Think before you speak: Training language models with pause tokens

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.376239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:06e79b4b043f9b172e446b23c2d215c65d62ce979afdeba8f7985be2f40dc968

Observation 448c6cfa-953e-4975-a68c-a7094cba5aa4 · outbound

This paper cites Generating Sequences With Recurrent Neural Networks.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Generating Sequences With Recurrent Neural Networks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.887084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:60ba27f700686ea938f02426bb9724f5189111bb7c4ea39ba74f19bb4cfa89f8

Observation a3b8d8a6-1cfa-49c3-8122-20fda793484b · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Adaptive Computation Time for Recurrent Neural Networks

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T11:53:22.734695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:bf67b981f37b4f9edda8a8490d928e245c234e292ca40411c982ffa6b5c75afc

Observation c93042f3-5792-45b5-bacd-e505f59884de · outbound

This paper cites Fast and expressive multi-token prediction with probabilistic circuits.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Fast and expressive multi-token prediction with probabilistic circuits

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.512028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:1cc9a35e027717743a3918dc6af8dd43bd6b82a32ecc161ddcff706e5c832e80

Observation df9c3301-d60b-4170-9348-944cce3b88c7 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mamba: Linear-time sequence modeling with selective state spaces

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.434981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:be093afd1cce2d3e66e2ebc5691cfa05cb751ff03fa92fa95ba36d1def2f0bbc

Observation e179736f-a2fd-408d-afde-1ea02f97e1af · outbound

This paper cites OLMES: A Standard for Language Model Evaluations.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models OLMES: A Standard for Language Model Evaluations

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.120033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:eff2f8821b068cdceaf93a1a05d65ab2e37c31da898f4e93df57a80b2ee9376a

Observation 52d0423a-5349-463c-b03d-a209ee79f5e3 · outbound

This paper cites DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T06:36:29.240841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:26a1321129912c3b17fd99616d700d4c7bc4f4daefe34fea5fc55c8bd5a28e06

Observation bd8a292b-ab5b-435c-970d-6f08198a24d0 · outbound

This paper cites General-purpose, long-context autoregressive modeling with perceiver AR.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models General-purpose, long-context autoregressive modeling with perceiver AR

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.428019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:86566e31e1f30d88119d3df997ee7671ba5abbc15c03ac61c0d97902eb0b154d

Observation 0d9e771a-d6fa-4a48-88ec-1e9de7d0b470 · outbound

This paper cites Measuring massive multitask language understanding.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Measuring massive multitask language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.420285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:24239abc03ea695bd62080ce8ef7204fc5ea995f15f9937d31a68106e63f1649

Observation 6fb07288-5514-4f0a-ab1c-cae81b06a61c · outbound

This paper cites Block transformer: Global-to-local language modeling for fast inference.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Block transformer: Global-to-local language modeling for fast inference

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.438923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:0c7aa22721498b984cec2f2a565f5b9d7a4518424d6d4537c34434bddd8215fb

Observation 763c9027-4771-4aaa-9e7a-596e99e52afd · outbound

This paper cites Deep networks with stochastic depth.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Deep networks with stochastic depth

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.416367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:84805725955c520545e208044bad95ea46daffe2e53cd4fa046549d8bae8134b

Observation 29771bf6-6e85-4dd7-b0a8-9a157412cf50 · outbound

This paper cites Conceptmoe: Adaptive token-to-concept compression for implicit compute allocation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Conceptmoe: Adaptive token-to-concept compression for implicit compute allocation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.009644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d757ec9b89ba5345241fe1581ac95dfa197fddaa512c75eed810e761c804c1fb

Observation c19a60f8-b258-42c8-85ac-9e512f693668 · outbound

This paper cites Character-level language modeling with hierarchical recurrent neural networks.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Character-level language modeling with hierarchical recurrent neural networks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.281143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:71d15ea61d1de81349891cf31aa3ea7c258c7f8504da7ec7ee28b6eda61c4840

Observation c9d4a8af-d64c-413b-9b18-497755608a41 · outbound

This paper cites Dynamic Chunking for End-to-End Hierarchical Sequence Modeling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:28.977825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:8c22a60be77ad3ec70c5d7e6c0a7ef040f2a2f648d90314c2ccbd7aa42beafe2

Observation 2fa510aa-69a7-4e8d-bc35-d13d9db76557 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:47:14.368858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:45cdab314c39b0abb06df4ac3dddecafc1e0444bb716878ed21bbd5d964acd29

Observation 32ec8fda-688a-48da-a1da-2688492f6782 · outbound

This paper cites Perceiver: General perception with iterative attention.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Perceiver: General perception with iterative attention

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.383706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:062678644029b1bcd11a44a4d2caf8c778700e4834a8923758ba45d98e2e6a67

Observation 303385c4-0d43-4ac6-b8f8-3d9d6249d87c · outbound

This paper cites `` low-resource '' text classification: A parameter-free classification method with compressors.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models `` low-resource '' text classification: A parameter-free classification method with compressors

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.343094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:9f0a551801047030072f54b0166dc69452ab41481831f86ab2aa61d604654573

Observation b180b5a6-dc70-4904-b5e7-e5b1062322d6 · outbound

This paper cites MrT5: Dynamic Token Merging for Efficient Byte-level Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MrT5: Dynamic Token Merging for Efficient Byte-level Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.928013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:40171e11833f62928ceede868bf3d03c9e7d431c64ecf53efcb61fb940e61f24

Observation abffccac-c31b-4d9e-902c-40b6ea87487d · outbound

This paper cites Subword regularization: Improving neural network translation models with multiple subword candidates.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Subword regularization: Improving neural network translation models with multiple subword candidates

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.353721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:768385859828a4ba30a0f5d8f136c1172d43b5ec5077186a9b2fac995da4ebbe

Observation 3d2df5aa-93d6-40e6-b645-f69dba7fb476 · outbound

This paper cites S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models S entence P iece: A simple and language independent subword tokenizer and detokenizer for neural text processing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.285555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4155ebab2d5b7649c187ac358f8e4bcce6dd3952823762166dfb5180c3259abb

Observation 9e563b4f-ee86-416c-bfb6-345a22a3786a · outbound

This paper cites Mamba-3: Improved sequence modeling using state space principles.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mamba-3: Improved sequence modeling using state space principles

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.399400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d6e5e08afc1c05d0c7270b3813f25a8cbe33378e4f4dacf8d8598ccde553dbbf

Observation 7861d0da-8c14-4247-a758-834f7c2ff845 · outbound

This paper cites Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.075646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:f5e2b72fc1577d32bfb94608191ec9d9acab169cbb3c0b87341273855c239b11

Observation 2f26fe72-8ceb-49ad-9377-9faedc558799 · outbound

This paper cites Training LLMs over Neurally Compressed Text.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Training LLMs over Neurally Compressed Text

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:28.952129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:773cd3869bd39db68fd09726346128c42f47b36d4376b4ee477f637937ef6467

Observation 1d5eba93-1f43-4c78-9e57-d6d325bc3670 · outbound

This paper cites DataComp-LM: In search of the next generation of training sets for language models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models DataComp-LM: In search of the next generation of training sets for language models

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:58:17.761776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:1d659be7feedd0525b229ba88a4e73d2a777fdadfa6d63121bdf1c8bf35ee321

Observation 9699f24d-0927-420a-a535-75180591fe3d · outbound

This paper cites MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MYTE: Morphology-Driven Byte Encoding for Better and Fairer Multilingual Language Modeling

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:28.968337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:33cd2cd7ff6bf2a4fe203e586449f6db72cdbcb7ff7ca849c4bce56785bdfc5a

Observation b91d14af-756d-4463-bb37-1c242002f7aa · outbound

This paper cites Smith, and Yejin Choi.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Smith, and Yejin Choi

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.277059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:b251e43371617f7ac072ef73c795a15e524dfa10ce3215538d50da2fc6d1cf4b

Observation 29744b20-b3d1-42b1-ae19-f3cb4d3bef1f · outbound

This paper cites Decoupled weight decay regularization.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Decoupled weight decay regularization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.350401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:eb0a7dd72aa11218cfa786827570dfffb043fa96222512181dc8247ca381636d

Observation b20de449-b3d5-4d72-8919-2c90cc6c6d94 · outbound

This paper cites Text rendering strategies for pixel language models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Text rendering strategies for pixel language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.405883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:3c332deb2dd45acb5dad7e150e79a85e417980341fd90c6498f9919430556030

Observation d928f038-711e-4ed1-b0cd-7249749f9f33 · outbound

This paper cites Starcoder 2 and the stack v2: The next generation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Starcoder 2 and the stack v2: The next generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.500264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:9e0a06ace8fd586473f109f9127f34ef46d9072983e177d764befef406e9e6ea

Observation 4ecb3e0d-da08-4a44-86df-edf53afc970e · outbound

This paper cites The art of prompt design: Prompt boundaries and token healing.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models The art of prompt design: Prompt boundaries and token healing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.503996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:b23ace4a12bb5aaecef23bd4f553ff590ff1e1a855131a60adcf5daf96888b8b

Observation 319fe42b-e335-4e72-b036-6550ee642834 · outbound

This paper cites Guidance.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Guidance

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.496641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4ee519cc70df0f33411151e70a5e71d6845dd317597d0b57e1b113333e55116f

Observation 323195a1-9d0d-4ac9-9322-5e67180d5615 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.328161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:ec1fb1202bd40f654ba22bbd7afefd683f4d0c7b72f768e6f063afc5dde68b97

Observation b6b19845-325a-49f3-846d-2831fa844e7b · outbound

This paper cites Minixhofer, T.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Minixhofer, T

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.206351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:beef9b75ac6e2900203bc9640a819af0aad4e8b3a067391589712724871ca55e

Observation f2e5a378-ef06-4093-8ff7-0a407bcee020 · outbound

This paper cites Hierarchical transformers are more efficient language models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical transformers are more efficient language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.450872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d0dc927870b020246492c9cc2865be5f9f65ae4502be0ab5b23bfdbb6d167be3

Observation 6c428cfe-862d-4929-a668-421fed0f9a22 · outbound

This paper cites Efficient transformers with dynamic token pooling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient transformers with dynamic token pooling

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.315000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:91fca2f1b00e1aa84c20f087a0fb3bf5d9c80042dafdd9e26ff2e110bfe34c31

Observation e2b73c34-d868-473f-b1c9-35637e98a049 · outbound

This paper cites Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Hierarchical Autoregressive Transformers: Combining Byte- and Word-Level Processing for Robust, Adaptable Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.879364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:6f5fbd3a26c8f2e3e0039a2160c2e39d07585c1a8775bf5348489b83c8a3a5ce

Observation 2a0d88a2-9ad4-4115-af90-bdca6c3c9326 · outbound

This paper cites GPT-4 Technical Report.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models GPT-4 Technical Report

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:36:28.914004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d0ea0c1a4cefeaea3ca33e44dce765c8759bd118d2e9971be6eb20d4147d65fe

Observation cd7fb21d-0d92-40c9-b092-663f8dacf775 · outbound

This paper cites FLEXITOKENS: Flexible Tokenization for Evolving Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models FLEXITOKENS: Flexible Tokenization for Evolving Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:56:13.953344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:43a1b4ff6f5c8c2abd8e83c14c716dfaf555a54d1fea5460971fbf59225d764f

Observation 53a577bf-b816-4a10-ad63-442b74c0309b · outbound

This paper cites Byte Latent Transformer: Patches Scale Better Than Tokens.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.184901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:6e8ab1d584698e65082fb2ec6d5b6134e3df97c0164d2347be0822f9e49f2818

Observation 626da5f5-2f83-4268-93c0-77ac5a6a2eaa · outbound

This paper cites Openwebmath: An open dataset of high-quality mathematical web text.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Openwebmath: An open dataset of high-quality mathematical web text

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.323155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:c928c49f6c2bac717eafda3e938f34ce1b7f7f5454b786884195a5aef76bb53e

Observation 5d331cc4-1b11-4e5f-be0f-46429cf8d53c · outbound

This paper cites Dynamic large concept models: Latent reasoning in an adaptive semantic space.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Dynamic large concept models: Latent reasoning in an adaptive semantic space

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.796718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:25fc2b98d05d339bc25a1a46db86353cde4f9ead57d269de0daf51bac3693379

Observation ea47df25-b79d-4ca1-9240-1bf4cbffcfb2 · outbound

This paper cites Learning to Generate Reviews and Discovering Sentiment.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Learning to Generate Reviews and Discovering Sentiment

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.039412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4297891a6ac397b81f8a20e151e5298b2578e83fd4384ff6221d8adb1a87d9b4

Observation 73f206ef-6d41-4de0-adb0-ac5abc801272 · outbound

This paper cites an unresolved cited work.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-12T16:46:42.412942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d35f17b7353ff1efaf5001abbbd2305c6c91ae94dc4cb63abdfe7c22adf3dabb

Observation 04407f9a-01d5-44cf-9868-0dd310ee8d78 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:02.305925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:ea1432bb394df01ccf4b362b3c4598b1b6a5cff7cef7a69f61c946e334fcd89f

Observation a776c116-46c5-412f-a994-dc32c9c80966 · outbound

This paper cites Solidgoldmagikarp (plus, prompt generation).

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Solidgoldmagikarp (plus, prompt generation)

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.454477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:92329c0963640ebd3737436ca2ee776285339c95a9e26826294982cbfda9e2e6

Observation e7f05fa1-038f-42ef-82d5-bf77bda3a08f · outbound

This paper cites Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Lotz, Emanuele Bugliarello, Elizabeth Salesky, Miryam de Lhoneux, and Desmond Elliott

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.465256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:b75f99705ed4e34dd63ced36899761dd0389bc68d2b67cc74ac11e07c26c00ec

Observation 75c4f69a-aa95-4c86-be27-ff4bfa0b4f73 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Winogrande: An adversarial winograd schema challenge at scale

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.368815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:7cf72dcb69f9b5bcaf6c882a695f6a0403d147afb0ff339eded94992b2aeb189

Observation 9c91c57b-26cb-46b0-b7b0-7aea5313010c · outbound

This paper cites Robust open-vocabulary translation from visual text representations.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Robust open-vocabulary translation from visual text representations

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.357496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:366ea343beb707fb94753ad53d3c470571b5b16864a876fc67deac07d8b92d88

Observation 1358df84-fcff-433a-8038-2887c466d418 · outbound

This paper cites Japanese and korean voice search.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Japanese and korean voice search

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.339496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:07f01b35233e7acb312947bb8c435872165da450f74788a8421c212d5ecb3f43

Observation 0cae3c0e-7e90-4dea-9d85-9c8524f753b2 · outbound

This paper cites Neural machine translation of rare words with subword units.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Neural machine translation of rare words with subword units

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.424027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:0a8f97a852c2eea9e63850b970cab198c0e55bbe79063664eb3c211b45934d8b

Observation 9b5ee4b2-cd9a-4ea0-b7ab-431894f91dc0 · outbound

This paper cites SpaceByte: Towards Deleting Tokenization from Large Language Modeling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models SpaceByte: Towards Deleting Tokenization from Large Language Modeling

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.090174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:8d698bbd832116895ef6c8a8f7ff7b2923feacf3568d48847478cce29eb182da

Observation cb630b05-c1e7-4b16-8bd7-fa16816d1380 · outbound

This paper cites Blockwise parallel decoding for deep autoregressive models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Blockwise parallel decoding for deep autoregressive models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.290011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:b5562702e8498545f112b38fc8869aa7f94ea64a6a4d8014d0697bcdd5e2e13b

Observation 4662c9dc-dd95-470e-bfde-e5aed21e6f98 · outbound

This paper cites Generating text with recurrent neural networks.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Generating text with recurrent neural networks

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.361809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:60b8000551095d50e604e22032685e7af0f81c7a7bf1f13c12b341c2d692af57

Observation 818ac257-47a3-4845-b669-46133384fa50 · outbound

This paper cites Sparse universal transformer.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Sparse universal transformer

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.431444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d517d4d05f695061d30faf67e970dc84290768cbb5f9958b762b6acadd213b9f

Observation 3fd762a8-b3b8-45ab-9edf-0c209165ee4d · outbound

This paper cites Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Tran, Sebastian Ruder, Jai Gupta, Hyung Won Chung, Dara Bahri, Zhen Qin, Simon Baumgartner, Cong Yu, and Donald Metzler

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.365289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:09a7ee5a837427692218d2e501369c1b290fcf9113a6a077f7929f34bc444cc6

Observation b906c358-94d1-46ea-9b78-cf7d5c7a9722 · outbound

This paper cites From Bytes to Ideas: Language Modeling with Autoregressive U-Nets.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models From Bytes to Ideas: Language Modeling with Autoregressive U-Nets

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:1508882af5181ffce3e9cb3db97c86790a71fd1a2da40de17f0355382ef55144

Observation eb950473-1012-45be-9eac-ab678a94bd2e · outbound

This paper cites MambaByte: Token-free Selective State Space Model.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MambaByte: Token-free Selective State Space Model

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.110367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:5d1d2636bb1b7ce9581b1baa9c7010a1c2d7bfbc6858e72c12c45bdf7a2b8c91

Observation 1aed69f3-a39b-4bf3-9295-2f7d1846d93f · outbound

This paper cites Parallel loop transformer for efficient test-time computation scaling.CoRR, abs/2510.24824.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Parallel loop transformer for efficient test-time computation scaling.CoRR, abs/2510.24824

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.026786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:e96b0f99507123c08d1ddda01b33fd47fe0e844a3616731daa4b7b3c983245a6

Observation f683faf2-3189-4eec-9029-366a612255d5 · outbound

This paper cites Efficient Pretraining Length Scaling.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient Pretraining Length Scaling

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.285778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:e4e3397dec1de6bd98aee43ce556f911726783e107db1ffbb19133b8b42a8b08

Observation 6c881321-048e-4011-9d33-d52fbfe35ff7 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:21:29.028955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:4222bcb599cfae18a8010903bca3c6463f87c4d04e692c3ee8ea8a576bcd2d74

Observation fc9d9a3b-a53c-445e-a988-fe3fd882444f · outbound

This paper cites B y T 5: Towards a token-free future with pre-trained byte-to-byte models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models B y T 5: Towards a token-free future with pre-trained byte-to-byte models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.302910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:6b4f536e7c243abdd3c6616402854c9083694c3090c97a627bf1f149cf13ffd7

Observation daccf9ec-9873-4498-94a3-b761c17004b8 · outbound

This paper cites Problematic Tokens: Tokenizer Bias in Large Language Models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Problematic Tokens: Tokenizer Bias in Large Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.172693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:dec466e1f459f870f71aea6174042af137c16d053194ebbcdd67a651a52e394e

Observation 77da302a-c137-4766-8f78-592c9548acb2 · outbound

This paper cites Scaling embedding layers in language models.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Scaling embedding layers in language models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.493134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:00d9955d9d4ef517be934512df1aa93877a2c3b76302f53b9d921d9033f8d83c

Observation 17cbaf07-d478-487a-9ddb-132efd34862c · outbound

This paper cites MEGABYTE : Predicting million-byte sequences with multiscale transformers.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models MEGABYTE : Predicting million-byte sequences with multiscale transformers

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.346658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d099d53d2db667000360da0aaa459fe2629cde2c6fe490c356e671c4419cebab

Observation ddd6ee95-0fdc-4e84-9b82-880d6749e8ac · outbound

This paper cites H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models H ella S wag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.335783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:2e25b8caee462c3a62ffcce2b7fc49de21989c89089ae6658f9ec99b260914c2

Observation 7d255cad-4e4f-4144-8e01-2a0a326d65e1 · outbound

This paper cites Ponder LM : Pretraining language models to ponder in continuous space.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Ponder LM : Pretraining language models to ponder in continuous space

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.481482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:455d81d26b7a56d1e1ec133e078f49695e3d235300a02f931c334dac5749765d

Observation f24ce0fc-075b-4fac-b0fc-10ba1da542c1 · outbound

This paper cites Linear complexity randomized self-attention mechanism.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Linear complexity randomized self-attention mechanism

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T16:46:42.372163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:d7974a5a4d2f87683515a00ad7a3681972dec01de2a91dd1a72f29217018f6d9

Pith citing papers

No inbound Pith citation observations are available.