Pith. sign in

Paper Citation Record · LEDGER

Byte Latent Transformer: Patches Scale Better Than Tokens

As of 20 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 44 inbound Pith citation observations for arXiv:2412.09871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09871 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:39.195688Z

measured 96 of 96 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:33:18.628297Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:47:48.461365Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d369c27-aa05-4d61-af2b-cfa6d1fbc0a6 · outbound

This paper cites write newline.

Byte Latent Transformer: Patches Scale Better Than Tokens write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.066549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.066549Z digest=sha256:084b4e53d5a8b5a2197ee3d822caa23abfeb906768597f5818a77b0992158440

Observation 693a9cad-1457-40e7-a424-220f9a8df050 · outbound

This paper cites Character-level language modeling with deeper self-attention.

Byte Latent Transformer: Patches Scale Better Than Tokens Character-level language modeling with deeper self-attention

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.576723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.070566Z digest=sha256:5103900859534d70e81e9b83224d3e9cfa3115baadbbd082005f04148fc8b922

Observation cba0e444-0c7b-49fd-9f14-e15af1736dcf · outbound

This paper cites Program synthesis with large language models, 2021.

Byte Latent Transformer: Patches Scale Better Than Tokens Program synthesis with large language models, 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.073285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.073285Z digest=sha256:8968b75fafa3c4cb1368751eba29a56c1b03bc7b4e0d1dba8a5abcdfdec0b7e1

Observation 90c30d31-d107-4ede-8a6b-85e9ee2efc52 · outbound

This paper cites Learning to rank with (a lot of) word features.

Byte Latent Transformer: Patches Scale Better Than Tokens Learning to rank with (a lot of) word features

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.564162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.076082Z digest=sha256:4e4b99b09493a007d9ebf1220c45647ddfca6b3e5dc66b48476a51118f5f393d

Observation cf52b182-05b7-4f9d-ad8b-e0ead61a17f0 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Byte Latent Transformer: Patches Scale Better Than Tokens Piqa: Reasoning about physical commonsense in natural language

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.556081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.078669Z digest=sha256:c8871b283f3ef000f734cf25184fdf000ecb09455e3348a04069518ab05c7876

Observation a8fe8fa5-940d-4194-9c0c-affa3b7d8ae8 · outbound

This paper cites Transformer flops, 2023.

Byte Latent Transformer: Patches Scale Better Than Tokens Transformer flops, 2023

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.548252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.081180Z digest=sha256:a96621cc69b1ab1ca3b32918d7f111494b9255e3521c60d7ada34238d54a148e

Observation 5b7afab3-5a82-4bb1-81d1-6e72a14c32a8 · outbound

This paper cites an unresolved cited work.

Byte Latent Transformer: Patches Scale Better Than Tokens Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.083672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.083672Z digest=sha256:c3cf32cb90c9bcc98c557069b1654fa2f79cdb94252bdfd9f3bc98c46a9418b7

Observation 9dc9a6f5-578e-466b-ac02-96b8bbde30ce · outbound

This paper cites Bridging the Gap for Tokenizer-Free Language Models.

Byte Latent Transformer: Patches Scale Better Than Tokens Bridging the Gap for Tokenizer-Free Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.086335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.086335Z digest=sha256:7ace5ca72a1351bfa673adfe5c202354646bd5edbe21ed6cd3d47b2512cf6ecf

Observation ce895d6c-f361-463f-bf89-011fb947120d · outbound

This paper cites Hierarchical multiscale recurrent neural networks.

Byte Latent Transformer: Patches Scale Better Than Tokens Hierarchical multiscale recurrent neural networks

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.533922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.089300Z digest=sha256:dc0c4110b9c6ab8db05bca64140a2d6baedbfce61f3f9f7123e255c26c76c735

Observation 4b746304-68f1-4a60-ae44-b8c384b897fe · outbound

This paper cites Canine: Pre-training an efficient tokenization-free encoder for language representation.

Byte Latent Transformer: Patches Scale Better Than Tokens Canine: Pre-training an efficient tokenization-free encoder for language representation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.525024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.091884Z digest=sha256:adab7a5f476f08f0633fbe04291198c40aaf6f01c32de19a2de4aad4d205fa01

Observation 0cdafcf7-2570-430d-b90b-f30140e021a5 · outbound

This paper cites Think you have solved question answering? T ry ARC , the AI2 reasoning challenge.

Byte Latent Transformer: Patches Scale Better Than Tokens Think you have solved question answering? T ry ARC , the AI2 reasoning challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.515856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.094435Z digest=sha256:7b80ce8375e347ce8eca944ecfcbae0fad1b3d5498a00c1bb5a62fa0ad840cf3

Observation 14bcd589-c480-4199-b336-ab787b060119 · outbound

This paper cites Getting the most out of your tokenizer for pre-training and domain adaptation.

Byte Latent Transformer: Patches Scale Better Than Tokens Getting the most out of your tokenizer for pre-training and domain adaptation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.507372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.097106Z digest=sha256:ebec03608b803a268c7f770b71372efe4cf545065635d3928fee281d78b309ce

Observation b386dd3b-498c-4a71-b9c8-81d30089216d · outbound

This paper cites Flash A ttention: Fast and memory-efficient exact attention with io-awareness.

Byte Latent Transformer: Patches Scale Better Than Tokens Flash A ttention: Fast and memory-efficient exact attention with io-awareness

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.499152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.099631Z digest=sha256:5564862c1b777455e27c7bc8fb132aa18eb36f3dab0a816a64506f305504b484

Observation 4fe1d1f3-2faa-48e1-a66e-9903e32e0482 · outbound

This paper cites The llama 3 herd of models.

Byte Latent Transformer: Patches Scale Better Than Tokens The llama 3 herd of models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.102056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.102056Z digest=sha256:42f4d45b351516eb8938ce441590d051f7e3b498aadfa696c69c54659a7fea4e

Observation b8f49ea3-17dc-469a-a24c-c31278ade781 · outbound

This paper cites CUTE : Measuring llms' understanding of their tokens.

Byte Latent Transformer: Patches Scale Better Than Tokens CUTE : Measuring llms' understanding of their tokens

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.487022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.104591Z digest=sha256:0e21633eb57643331a1199c9cae80d7ee15f35393f3534317368084265e65035

Observation b5acf611-f4dc-492a-94e7-45bb832d4937 · outbound

This paper cites CharacterBERT : Reconciling elmo and bert for word-level open-vocabulary representations from characters.

Byte Latent Transformer: Patches Scale Better Than Tokens CharacterBERT : Reconciling elmo and bert for word-level open-vocabulary representations from characters

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.479824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.107175Z digest=sha256:9e12d0c5dd91f9311ff53d097486a0a5d0fe1e5637aa3a8d4cd875a5b12df0bd

Observation 87143f23-eccb-4df1-88fb-bde3c2bab50c · outbound

This paper cites A new algorithm for data compression.

Byte Latent Transformer: Patches Scale Better Than Tokens A new algorithm for data compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.109546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.109546Z digest=sha256:9ec33490511a8625d1d5a0bdda3f10470acee684ec131fe1f18cea108b1875bf

Observation d71886da-cd17-4312-820e-719cd3150153 · outbound

This paper cites The F lores-101 evaluation benchmark for low-resource and multilingual machine translation.

Byte Latent Transformer: Patches Scale Better Than Tokens The F lores-101 evaluation benchmark for low-resource and multilingual machine translation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.112003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.112003Z digest=sha256:b7ac2e476782a9392675fb88da1c96126944b882ec9fc12f3745479a76aa1bf6

Observation db7ffea0-0c71-4771-85c2-a2db1606698b · outbound

This paper cites Generating sequences with recurrent neural networks.

Byte Latent Transformer: Patches Scale Better Than Tokens Generating sequences with recurrent neural networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.468438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.114610Z digest=sha256:c6eebb8025e32f1390f53e872cd670129db6a9ba569eb7fa9ab390db02e230a4

Observation 5aa62707-6ff5-45ed-9ee7-81a236097e48 · outbound

This paper cites Mamba: Linear-time sequence modeling with selective state spaces.

Byte Latent Transformer: Patches Scale Better Than Tokens Mamba: Linear-time sequence modeling with selective state spaces

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.461482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.117348Z digest=sha256:151cb3b437a5ad0b59a8a30df7bf49ba137d2ac750acf61b08e361eb5e889fb3

Observation 72238b77-86f6-45d8-9856-b16fe44bfff0 · outbound

This paper cites Measuring massive multitask language understanding.

Byte Latent Transformer: Patches Scale Better Than Tokens Measuring massive multitask language understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.454497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.119792Z digest=sha256:308c81ac2c53f1cc7d20472bb7b4a12ad8c5b61621cfa075b052df973fd000eb

Observation 380e68d6-1f27-4146-87f7-dfdc9a458a85 · outbound

This paper cites Training compute-optimal large language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Training compute-optimal large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.446929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.122183Z digest=sha256:fe00513701e3fb44e777e84e7c8e99264b245c542f51a4fb4e2ee45f6c411243

Observation 5f0a86ce-4196-445b-8f88-ee07e797f445 · outbound

This paper cites Perceiver: General perception with iterative attention.

Byte Latent Transformer: Patches Scale Better Than Tokens Perceiver: General perception with iterative attention

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.439623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.124668Z digest=sha256:0cec63416a2c1a8360d4c71409acc5db61c86271e8b0007c6ddd7ef73da25cce

Observation 3c809974-f515-4f4d-8744-5d2a407f73f5 · outbound

This paper cites Neural machine translation in linear time.

Byte Latent Transformer: Patches Scale Better Than Tokens Neural machine translation in linear time

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.432361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.127112Z digest=sha256:8d8199a716a4ef204349da70e22fb92468237dbd2edc1dee3d0bc1d8d9adb847

Observation 967658ef-d4ac-4249-af3a-a6d12550dd46 · outbound

This paper cites Scaling laws for neural language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Scaling laws for neural language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.425103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.129455Z digest=sha256:1551f74624542d440dc255385fa9deba787d00080b2015a30c3bf6e6cc319828

Observation 7e24b8f0-ff74-4360-8cf5-7b865c2555ea · outbound

This paper cites Byte-level machine reading across morphologically varied languages.

Byte Latent Transformer: Patches Scale Better Than Tokens Byte-level machine reading across morphologically varied languages

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.417857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.132151Z digest=sha256:f19a9bb790d3ee8b7b502e27b949f0ccc35c5f91b34c0c73a3690b513aadfee2

Observation b0f522bc-7e01-4e2f-b88b-5a0b38957714 · outbound

This paper cites Character-aware neural language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Character-aware neural language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.410015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.134608Z digest=sha256:f1461fe441d29c10e9f97a6066f438b98cc2635fc8a1302406a3be4133ce281e

Observation 730d1e14-8232-447a-bc9a-711598e7c305 · outbound

This paper cites Training llms over neurally compressed text.

Byte Latent Transformer: Patches Scale Better Than Tokens Training llms over neurally compressed text

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.402581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.137715Z digest=sha256:71a64cc811868f126c74a553f3a84e30309019757a1ff2ebe95e83efde05bfd6

Observation d2ad2aee-659e-4ee8-939f-55047a6d918b · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Datacomp-lm: In search of the next generation of training sets for language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.395059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.140167Z digest=sha256:a6505ce5d7e8335b9a7b51489da13582a667e39add0aa32459a8fc76a1fca78a

Observation f9ff057d-ea1f-4852-b6a7-90c101f4d4dc · outbound

This paper cites Xlm-v: Overcoming the vocabulary bottleneck in multilingual masked language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Xlm-v: Overcoming the vocabulary bottleneck in multilingual masked language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.387107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.142636Z digest=sha256:172cc13453c3b0c957b762d37b012398d24c13a4e98cfc9b1882d0dd3c8b5dbe

Observation 2293348e-1ad8-4bbd-8279-5515acc108ed · outbound

This paper cites Myte: Morphology-driven byte encoding for better and fairer multilingual language modeling.

Byte Latent Transformer: Patches Scale Better Than Tokens Myte: Morphology-driven byte encoding for better and fairer multilingual language modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.379218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.145173Z digest=sha256:d95b6fe5a020597d9b36eb1f74b2c7b6cad487b3fc2ca7bca73541a39d44dd52

Observation 7f0076ab-9dc7-4f70-b862-4e8ee5c97ee2 · outbound

This paper cites Decoupled weight decay regularization.

Byte Latent Transformer: Patches Scale Better Than Tokens Decoupled weight decay regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.147564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.147564Z digest=sha256:175b364ed8de58c5e9271e6ec739d911818efe6b1b23b33e1b6e7a9d746e2be4

Observation 9fa05016-eeb1-4dca-b62f-f56c836db36c · outbound

This paper cites Subword language modeling with neural networks.

Byte Latent Transformer: Patches Scale Better Than Tokens Subword language modeling with neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.367133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.150084Z digest=sha256:c34a6b0600240e1d8fa5060385beed9067b112752a06d7d01ec9e6e3ac49e78b

Observation fbb644b8-3b75-4b2a-ab6f-c2cc47dc369a · outbound

This paper cites Hierarchical transformers are more efficient language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Hierarchical transformers are more efficient language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.359075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.152669Z digest=sha256:bbf1595e16182c26a37682e2367c72272f3aa166490f5a2bd69d67a7a678b103

Observation 4fe4ac8f-66a5-4b9c-93c9-a9984becc668 · outbound

This paper cites Efficient transformers with dynamic token pooling.

Byte Latent Transformer: Patches Scale Better Than Tokens Efficient transformers with dynamic token pooling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.351233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.155111Z digest=sha256:342a0be250414c3cfc76693d2099699cd9dd2c00fd23ceb4b732246a08897b69

Observation 485d1b9b-f9a8-4122-bcfd-8a7eb38910f2 · outbound

This paper cites Language model tokenizers introduce unfairness between languages.

Byte Latent Transformer: Patches Scale Better Than Tokens Language model tokenizers introduce unfairness between languages

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.343278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.157626Z digest=sha256:24162c1c253925a977bf02881d0715cf0a0f89768c7ce19ab95c8411915cf566

Observation ca3519ec-3e35-427d-934c-578e910b3fb3 · outbound

This paper cites Language models are unsupervised multitask learners.

Byte Latent Transformer: Patches Scale Better Than Tokens Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.160221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.160221Z digest=sha256:afc3b32b13e5add1b9116e2426a51ceec8924a5b9282c3ca2e77c5edffbc93eb

Observation cbcbaada-7741-4cfa-8d9b-b774b8062ec3 · outbound

This paper cites Neural machine translation of rare words with subword units.

Byte Latent Transformer: Patches Scale Better Than Tokens Neural machine translation of rare words with subword units

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.331347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.162593Z digest=sha256:2840c0af90ba1be5e181269b34d98b760a15e3f69bbc3f442ed37a2003a991d8

Observation 510b83f9-51a5-4cda-8cda-0c8543e7a90a · outbound

This paper cites GLU variants improve transformer.

Byte Latent Transformer: Patches Scale Better Than Tokens GLU variants improve transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.323938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.165075Z digest=sha256:266ef358ea741a4d144862aef598cc94ffe8e0b66c70e43fcacd39df03334226

Observation e1a77165-4065-4ff2-baaf-771e5db04329 · outbound

This paper cites Spacebyte: Towards deleting tokenization from large language modeling.

Byte Latent Transformer: Patches Scale Better Than Tokens Spacebyte: Towards deleting tokenization from large language modeling

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.316186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.167400Z digest=sha256:b345f531e9483f4bd6a605801f0b3641090b81512ad0bffa417c6d08b5b3fb89

Observation 17801808-54a4-4813-85f8-185c26876ab7 · outbound

This paper cites RoFormer : Enhanced transformer with rotary position embedding.

Byte Latent Transformer: Patches Scale Better Than Tokens RoFormer : Enhanced transformer with rotary position embedding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.308897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.169797Z digest=sha256:47b330ee633cec8de000ee686b1f9aaf51c061349f0bcbe1bff006fc6607d839

Observation caf19b92-2190-4463-b62f-5ade5d799632 · outbound

This paper cites Generating text with recurrent neural networks.

Byte Latent Transformer: Patches Scale Better Than Tokens Generating text with recurrent neural networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.301286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.172076Z digest=sha256:dabdc1bd6cd241c3e9e1e67f525b2fd92b94122fb198e3b045756772d073e253

Observation 270c9470-2275-4a8f-b6d5-999758b15a60 · outbound

This paper cites Phonologybench: Evaluating phonological skills of large language models.

Byte Latent Transformer: Patches Scale Better Than Tokens Phonologybench: Evaluating phonological skills of large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.293961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.174368Z digest=sha256:8eea5a605e5d8818a73a7078b025b76f2cb64f8adede462140cf6d0457248f59

Observation 3ca97346-4abe-4d7b-8636-46d0eecc1035 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

Byte Latent Transformer: Patches Scale Better Than Tokens Llama 2: Open foundation and fine-tuned chat models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.286302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.176790Z digest=sha256:8a9be60f9a8e1d3b46b50155ef622bd29701f3974c1a2e66b2cde43e7148ddf2

Observation 5a9268c6-dc49-4989-a1ab-829d708d70fc · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Byte Latent Transformer: Patches Scale Better Than Tokens Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.179025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.179025Z digest=sha256:6faf9c16d934121c42ebcce9e819674b218763d8de644872b0214c4b663bb36a

Observation d045f22b-b877-4600-8092-1d1e512f0f5f · outbound

This paper cites Mambabyte: Token-free selective state space model.

Byte Latent Transformer: Patches Scale Better Than Tokens Mambabyte: Token-free selective state space model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.274226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.181129Z digest=sha256:7b7014f0f8f505dc3d4586e22eabde75528e79a4969eaae8a8956cc2fb6613ac

Observation d0a1bb76-62f3-451a-ab6a-9d47929cc9a8 · outbound

This paper cites Effective long-context scaling of foundation models.

Byte Latent Transformer: Patches Scale Better Than Tokens Effective long-context scaling of foundation models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.266582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.183492Z digest=sha256:b552642c0e70cb15fc9bee00ccfc7623c0dbdb5f9e6ca44eb23b6c1d5beca884

Observation 5a8c4775-7903-4985-a682-18a2366f5f28 · outbound

This paper cites Byt5: Towards a token-free future with pre-trained byte-to-byte models.

Byte Latent Transformer: Patches Scale Better Than Tokens Byt5: Towards a token-free future with pre-trained byte-to-byte models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:39.185967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:39.185967Z digest=sha256:008d22f38ce77408c58617cb0a9462bd43120fa80b2fa8b73f33c22afc2d44f6

Observation 48951ae3-14b0-4eb1-b4c8-d616b85b24eb · outbound

This paper cites Megabyte: Predicting million-byte sequences with multiscale transformers.

Byte Latent Transformer: Patches Scale Better Than Tokens Megabyte: Predicting million-byte sequences with multiscale transformers

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.254543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.188323Z digest=sha256:8bb6121443813ee7fff328bb5b73861e176f40e3d3aff20d99e33c26389aef6c

Observation a695aa24-681d-451a-8939-add0cd7362bf · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? arXiv, 2019.

Byte Latent Transformer: Patches Scale Better Than Tokens Hellaswag: Can a machine really finish your sentence? arXiv, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.246521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.190818Z digest=sha256:73e1857e0fc32b26b8ba3e442e85e604afc1cfc116166b27d41dd5208b48e98e

Observation fdbcdc63-6a07-4fc5-8fe5-83cbfe55d578 · outbound

This paper cites Root mean square layer normalization.

Byte Latent Transformer: Patches Scale Better Than Tokens Root mean square layer normalization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.237864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.193182Z digest=sha256:6c60d15ade289a4ceadb4b877616c24c9b7375d8a7369165904e853264407840

Observation 7d0a80c5-90c5-4e1b-9c44-a43cfab602d5 · outbound

This paper cites Character-level convolutional networks for text classification.

Byte Latent Transformer: Patches Scale Better Than Tokens Character-level convolutional networks for text classification

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:39.229061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T16:45:39.195688Z digest=sha256:c85beabc8cd6691e98384ff4a084d337e01bd3922569a24ae16e8fc1032f6506

Pith citing papers

Observation 8a56eef1-ae4b-46b4-907c-c68d04da976d · inbound

Tokenisation is NP-Complete cites this paper.

Tokenisation is NP-Complete Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:45:54.470359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:45:54.470359Z digest=sha256:d2f91de3fbe8780973116bdf1b52f10b90b7a4ef8669b25d711c1f3cdcdb6a8d

Observation 4326a6ca-de80-49ce-a5ac-fd0fff6d75e1 · inbound

Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers cites this paper.

Graph-Aware Isomorphic Attention for Adaptive Dynamics in Transformers Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:22.531753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:22.531753Z digest=sha256:5671980e2dfd57a4938bf6d6a35c366a2d2fba24340d13913c9932471c4c8b5d

Observation 126eaf8b-4455-492c-9284-6290608b28d0 · inbound

Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling cites this paper.

Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T05:28:04.841086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T05:28:04.841086Z digest=sha256:29dd6e0149f650604b7da72ff5a033bc0c923ad80ca96eb58cc5c9dacab33690

Observation 3d88bde0-ea82-48b9-b831-7338928bc81f · inbound

LLM-based event log analysis techniques: A survey cites this paper.

LLM-based event log analysis techniques: A survey Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T18:11:46.540034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:11:46.540034Z digest=sha256:9c9a60964bf427a97e3f3d546c0d5894f653de99f6d851a6c8057d2363be62af

Observation c8b92a72-72c2-4424-bec4-17a25157877c · inbound

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning cites this paper.

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:22:33.078204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:22:33.078204Z digest=sha256:69776a5e115fbc8b73d23c5bf49ed88847abf467ba36cbafce876426b2dca8b8

Observation 5bb9f469-f237-4676-8224-e6cc4e59f5f3 · inbound

Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning cites this paper.

Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T10:16:49.163383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:16:49.163383Z digest=sha256:8cd406317a0e09135dc76019ff90e582f2b75b0a275f8d9a1ccdb0e78d7ae158

Observation 21f1b5c5-f89a-4c3c-a615-48294aef94ca · inbound

An Uncertainty Principle for Linear Recurrent Neural Networks cites this paper.

An Uncertainty Principle for Linear Recurrent Neural Networks Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:15.898992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:15.898992Z digest=sha256:42df712ef0cb60bea26979ab19aa6348de6f4fd861c80c7650db328d2be99c08

Observation ca9a8266-a026-4ed1-be69-2768a92a6298 · inbound

ALFEE: Adaptive Large Foundation Model for EEG Representation cites this paper.

ALFEE: Adaptive Large Foundation Model for EEG Representation Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:33:18.628297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:33:18.628297Z digest=sha256:2e0c3ac215ed8f7a5c312fa684166c4b4b7e4937c2849230ee35f808e6f2673a

Observation 5a162554-a44c-43b1-b0a0-d211fd6668d2 · inbound

FreeMesh: Boosting Mesh Generation with Coordinates Merging cites this paper.

FreeMesh: Boosting Mesh Generation with Coordinates Merging Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:17.509395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:17.509395Z digest=sha256:817244f6e2bd9e43fedfa814dfdecbe109a56818c4a84be5ea38f28950803fe8

Observation 4f9c47a7-419b-464e-a43d-3d6ef8f0130a · inbound

EXECUTE: A Multilingual Benchmark for LLM Token Understanding cites this paper.

EXECUTE: A Multilingual Benchmark for LLM Token Understanding Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:32.255302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:46:32.255302Z digest=sha256:c12d1e703b39257446bc7d81eab8f9cfde5ff354dc3694e287793dd032333117

Observation 48674664-e288-4d0f-a3a2-a70e1a776f46 · inbound

Improving Language and Modality Transfer in Translation by Character-level Modeling cites this paper.

Improving Language and Modality Transfer in Translation by Character-level Modeling Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.989485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.989485Z digest=sha256:eb98d08d15fbd3c8e5beda1864e0b2621f545c3f16cdfe2542d6981c7e4b9b97

Observation 4a1085e6-f3aa-4d90-8b7f-dc8d684d1442 · inbound

E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models cites this paper.

E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:02.623401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:02.623401Z digest=sha256:425aeeb9fbe19316e877026275e9522021ae16ee2775a10443d8ca78e892b735

Observation 435b1c62-d240-45d8-be4e-5163e52aa492 · inbound

Sampling from Your Language Model One Byte at a Time cites this paper.

Sampling from Your Language Model One Byte at a Time Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:53:02.099512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T09:52:20.923710Z digest=sha256:3947afe3eca730bd5c67872c2d1eac9f56b867cbb33be60e158b8be06905acfd

Observation 13953cd7-5da7-420e-bdb4-4ca577edf254 · inbound

From Bytes to Ideas: Language Modeling with Autoregressive U-Nets cites this paper.

From Bytes to Ideas: Language Modeling with Autoregressive U-Nets Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T19:56:40.676381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:56:40.676381Z digest=sha256:e1c940aa35ecb4f604c24a7fc8dc891d0df9bebe8f3edac45db8386b37be7009

Observation 42ade3c6-115f-48e0-99a6-2e3a237063e1 · inbound

Entropy-Driven Pre-Tokenization for Byte-Pair Encoding cites this paper.

Entropy-Driven Pre-Tokenization for Byte-Pair Encoding Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-15T19:34:23.741958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:34:23.741958Z digest=sha256:cf0d95feec9f3a81c1287419621b491dd99a127850ef9d28757e75244f787fac

Observation 1db18ae7-1465-44d1-bff7-ba515ca024bd · inbound

ByteSpan: Information-Driven Subword Tokenisation cites this paper.

ByteSpan: Information-Driven Subword Tokenisation Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:54.053549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:54.053549Z digest=sha256:0b9b455e95d6d04e01d40e03ebc08fdc29b81654de7dd6ef1408f8f0ea76e774

Observation 2a6722b8-8bcf-42d5-ba56-645dc9e209fc · inbound

Learning to Skip the Middle Layers of Transformers cites this paper.

Learning to Skip the Middle Layers of Transformers Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:55.730794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:40:55.730794Z digest=sha256:fdf880049e67e55da04b120aee66ff6e4a2763eff8ede47ebfce5be219d705eb

Observation baca7f72-cda5-452d-9c71-b893ffbde6cf · inbound

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead cites this paper.

Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 280

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:46.494957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:46.494957Z digest=sha256:9b50e39c058580a914017ceef6b08e0e8fb341aa3963bf882dcf9097a42191b3

Observation 62add196-6752-4877-bae9-35fb7daa884e · inbound

Energy-Based Transformers are Scalable Learners and Thinkers cites this paper.

Energy-Based Transformers are Scalable Learners and Thinkers Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T20:42:36.683346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:42:36.683346Z digest=sha256:0555e55fe1c9fc4a5facd7c6f9e14b48034c10ba5fa59252aeb1ad23cc17df0c

Observation 8e67bcc0-2611-495d-9d96-e7291394e431 · inbound

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling cites this paper.

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T18:34:02.209101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:34:02.209101Z digest=sha256:aec53349e691a60af286b2fd601eaedc90e32826122d3e4afdc21fb398ad42bd

Observation 13ab9f17-58fd-4f22-91c6-1c85f375d120 · inbound

FLEXITOKENS: Flexible Tokenization for Evolving Language Models cites this paper.

FLEXITOKENS: Flexible Tokenization for Evolving Language Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:12:05.199207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T05:10:19.093601Z digest=sha256:3b19de9dc9f63e1225ffaaa06fdc85241d50a915393f62dc9ff6a2e377a4d4e6

Observation b6c96cd5-26b4-46c5-a8f7-55e1c4d53548 · inbound

Synergy: End-to-end Concept Model cites this paper.

Synergy: End-to-end Concept Model Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:34.701858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:45:34.701858Z digest=sha256:427891275b41a2b4231bdd29b3e24254d8a18c5c72c1a3560513045e71633e3b

Observation 245cef04-c4a6-4041-b171-e7fca218e328 · inbound

SpeLLM: Character-Level Multi-Head Decoding cites this paper.

SpeLLM: Character-Level Multi-Head Decoding Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:16.943166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:16.943166Z digest=sha256:67cb12da9686efab7a1865dbd0cd1fca6a5b7a0f31860efc2543f5ee796f47ee

Observation 38465e18-0e65-420d-a0d2-dcc00bdde284 · inbound

Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music cites this paper.

Amadeus: Autoregressive Model with Bidirectional Attribute Modelling for Symbolic Music Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:36.545581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:36.545581Z digest=sha256:057d59313fd536cde1506a8deaf40e5cf99250f4370b8a398ff67c53f7bac8fc

Observation a868451e-b1a5-421b-aa4e-550ebee901fb · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:21:15.070824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:26b26f8b5d1fed060821cd2ab62212669695ef8b458f21cf17045faae32a0140

Observation 800d5deb-5edf-4a87-b26e-47a6a5381726 · inbound

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking cites this paper.

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T15:10:05.833399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T15:10:00.146694Z digest=sha256:687979a277923917f412ee4c8e550b371d14a56a1736b7ed04d73e88461eb3f9

Observation 91b7b0a1-0333-4256-9c9d-5b51041eeb5e · inbound

Lost in State Space: Probing Frozen Mamba Representations cites this paper.

Lost in State Space: Probing Frozen Mamba Representations Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:13.005369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-09T19:57:22.493681Z digest=sha256:2cc00f07b6c5db8fb5161db71f88e42e535f19d5258f9a6e9bbc723f7e8cb85b

Observation 53a577bf-b816-4a10-ad63-442b74c0309b · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:36:29.184901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:1748273dd03934c38962479d263ece6eef92dc6089003ba75d6a3b71601751f4

Observation 7df2ea19-0fd7-49aa-9205-13bf95949310 · inbound

Bin Latent Transformer (BiLT): A shift-invariant autoencoder for calibration-free spectral unmixing of turbid media cites this paper.

Bin Latent Transformer (BiLT): A shift-invariant autoencoder for calibration-free spectral unmixing of turbid media Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:17.583602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T05:10:58.035490Z digest=sha256:bcbba4d7e412141d4f36624deeaf079de8f793f0c6a853a51ab794a9b06d5dc2

Observation b84543b0-8be2-479c-9859-cf5a6484eceb · inbound

Towards Understanding Self-Pretraining for Sequence Classification cites this paper.

Towards Understanding Self-Pretraining for Sequence Classification Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:33:58.896749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T05:29:58.809024Z digest=sha256:1385c1f2e42dc0ea6d71a22120a67e0e6b5958385d260daa1f8ce6ae267d982c

Observation ba6156fc-1975-4c98-86fe-078c7e3ebb5c · inbound

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models cites this paper.

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.401957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-29T07:50:25.019889Z digest=sha256:7ae354dba2698bacd0601d29fb774e22f75521b899667a410c9d62ae8bc97c82

Observation 3c62183c-42a5-4164-9825-5951dd2d4e5b · inbound

Large Byte Model: Teaching Language Models About Compiled Code cites this paper.

Large Byte Model: Teaching Language Models About Compiled Code Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:56:24.019070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T13:47:15.984105Z digest=sha256:a8ff15ba5000bcb1a8cc70cb0943f8ff689d001103a9e2f0b8a4fbe5198d1a93

Observation beafb45c-a498-40d1-a858-65063d412e83 · inbound

MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments cites this paper.

MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T04:26:35.422286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T09:13:26.501298Z digest=sha256:d09ece61771690dbbd7368384ae31c1b080951cce0fe2327b37b5ccd9ffb1408

Observation 9283bf06-bb4d-4426-a496-a5bb87b54656 · inbound

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models cites this paper.

Beyond Perplexity: UTF-8 Validity in Byte-aware Language Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-27T05:20:35.905563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T05:15:07.811143Z digest=sha256:c074a79b9743b2e1e53c0c0e41caa98e6d92fa30d18ec0f7084c2d41dd4847f7

Observation 22acbaf7-26c8-4e4d-93bd-f5e7789f4c1e · inbound

User as Engram: Internalizing Per-User Memory as Local Parametric Edits cites this paper.

User as Engram: Internalizing Per-User Memory as Local Parametric Edits Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:09:19.466162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T20:37:01.382431Z digest=sha256:5124fdfbac494d013969adae5a8c21099fab3e5c9856de3392118dc7b15dedf8

Observation 3b81a4b1-4169-4196-acf0-2eccc0eae93e · inbound

Protocol-Aware Tokenization and Architecture Co-Design for Wireless Packet Foundation Models cites this paper.

Protocol-Aware Tokenization and Architecture Co-Design for Wireless Packet Foundation Models Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T19:25:00.475146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T19:23:59.334530Z digest=sha256:952e04cfdfdeda692089bdd31417c09b9b21f2be4a8e56e50352f29615768e48

Observation 9cc5f5e6-a47f-4b00-8cf0-083dd64b678e · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.623582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:d387dc7d14f21fd6f44873dd80b289626fc0b5e7f4ec1df4176f7cf2aa847fc0

Observation 2cdd0554-5906-4ba3-933a-f63d452445c0 · inbound

EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction cites this paper.

EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Prediction Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:35:59.085284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-29T01:51:42.113800Z digest=sha256:7ba78c4cbc8b47919c7b743d1f33019ca45f20931cbcf61ee58996da97f6851c

Observation 13391aad-58d9-4548-a284-6bc4e487ded6 · inbound

Cybersecurity is the True Frontier for Generative AI Success or Failure cites this paper.

Cybersecurity is the True Frontier for Generative AI Success or Failure Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:54:35.389902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T09:46:06.726675Z digest=sha256:5489bdd641b22d65f6fd8456deb3f51c2b8d916f9efc43b74d2edb9d97dc4f73

Observation c451fb12-4662-40ad-9d8e-929f82ac05e3 · inbound

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking cites this paper.

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:28:18.347271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-03T13:20:21.237349Z digest=sha256:e6b514f7458a610834302cf8188a583dcc518df5a9e640fbd670bf482122c1dd

Observation 8755bc1c-bb44-4e9b-a4f2-7bf298987e22 · inbound

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking cites this paper.

SUNTA: Hierarchical Video Prediction with Surprise-based Chunking Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T16:42:04.979573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:42:04.979573Z digest=sha256:e76fe2a9719212cb9522ed23b8de9ace5645c07a74cacaccf031142b6131f179

Observation 05d66f14-9838-4e25-9a40-c438a24c5c58 · inbound

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES cites this paper.

Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:47:48.485819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-11T03:42:21.307552Z digest=sha256:ba80e1f5ab37293f3c3deb604df4f21640ebcff679c9b377726ddbc786701e91

Observation 0799d3ef-482e-4656-ad44-e5616d0dca69 · inbound

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization cites this paper.

Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T05:09:56.046733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:09:56.046733Z digest=sha256:9de757b4aeaf2873a82fd7ea16d57a02e28198cfdd22adbed74e0c3740d792e3

Observation 9d6c47de-2dd6-444f-a2bb-d543f28ab848 · inbound

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness cites this paper.

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness Byte Latent Transformer: Patches Scale Better Than Tokens

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:27.271301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:27.271301Z digest=sha256:25e20e372c916b5ce5e602fd96da0a84f4d3b82bd879aa25ffbbe68c982da2fe