Pith. sign in

Paper Citation Record · LEDGER

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

As of 20 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2505.09738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09738 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:30:03.795718Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:25:13.091782Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:38:54.307688Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e2aa0f6-8172-4036-88b6-6c6167d80146 · outbound

This paper cites ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.740361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.740361Z digest=sha256:0f667a4a40dd73bb00ec9be8cfddc82f618e6e6be39e85315c95567ec216d85a

Observation 22befeec-313f-4062-ac01-362ad341eb28 · outbound

This paper cites An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.748757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.748757Z digest=sha256:1ce7ac63a2337492d8caa511f338e35d78890a816463173b077f09e1c9c345f6

Observation 73b21c49-599b-4ed1-a22f-805c7da9fc19 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.753278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.753278Z digest=sha256:6a6ec7990a41e35151499dc8726695488a852ede79d72c5b0aeda52295b013d8

Observation dd926726-e4b4-429c-bd75-a986b2989689 · outbound

This paper cites Airavata: Introducing Hindi Instruction-tuned LLM.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Airavata: Introducing Hindi Instruction-tuned LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.758013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.758013Z digest=sha256:2bac95b949092fdd4e0cff44748f5b8a96097b63bc76dee42f2e80a0ccb4b9ab

Observation e6088e76-f291-4698-8aa2-21c1e9409625 · outbound

This paper cites Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.762381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.762381Z digest=sha256:2b49bf94c8ef57b6f38e320fa1ef54ecff3f0bdf24402a96c23bfacd715a3826

Observation 6239edb4-55f0-4240-ae0b-84260c0ed6d1 · outbound

This paper cites WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.774027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.774027Z digest=sha256:6f6285902f4b61dd6772746428688e9fc86e6e9cf46de1bb29eade420cb7877f

Observation 8ac63071-f400-43c9-9dec-3435350947f5 · outbound

This paper cites Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.777814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.777814Z digest=sha256:3b21eb1db4e0c96ff8ca3c8b921a09a6e847ade0ebedf35ca15eda635b067f46

Observation be618465-4cbb-41db-b90e-2ecc74ea96af · outbound

This paper cites The ultraspherical rectangular collocation method and its convergence.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning The ultraspherical rectangular collocation method and its convergence

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:30:03.877668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:30:03.786577Z digest=sha256:d0486f77ffa7b6b7ccf42dd786af7c1e3e644b491aed98188564e0d09d21b33b

Observation b0a10a52-e42c-40da-95eb-fcbe414bd2c3 · outbound

This paper cites Attention Is All You Need.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Attention Is All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.791522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.791522Z digest=sha256:9acba41c3e0e2c4ee2b938458651f295eb39f06253cc2fd820bc372a3e31f5a6

Observation 9b6c678e-1e54-4153-b26a-4cde16206477 · outbound

This paper cites Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.795718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.795718Z digest=sha256:6129e7c9a8f431b766d586f7cbcd8f839a7262a979652a0827cc325a43ae3570

Observation 315fa27d-32fe-4ab9-88b4-804ca606584a · outbound

This paper cites doi: 10.18653/v1/ P16-1162.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning doi: 10.18653/v1/ P16-1162

Reference 2016

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:30:03.782214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.782214Z digest=sha256:67e6c8a139c74067be6bb7337955d7e4ebf2bf2ccd4a3844c254b78c8e9ad01c

Observation 67aba5da-2eeb-4886-8c77-8bd1c2247c84 · outbound

This paper cites an unresolved cited work.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:30:04.015946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:30:03.766499Z digest=sha256:063bc73ce4ab838a7a1091700f18258505de2619c79e86af7fdeb263624a19d1

Observation 644ef3b4-a3a6-4d1c-9612-bfde45b66f07 · outbound

This paper cites Tamil-Llama: A New Tamil Language Model Based on Llama 2.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Tamil-Llama: A New Tamil Language Model Based on Llama 2

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.736366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.736366Z digest=sha256:8b963645b161d60f87143a9ef0c45c0e2bd00a72ac322ac84b987dbb4b1777ee

Observation 07b60db6-c8cc-41de-b76e-b130947d0468 · outbound

This paper cites Entanglement generation in capacitively coupled Transmon-cavity system.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Entanglement generation in capacitively coupled Transmon-cavity system

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:30:04.003902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T21:30:03.731822Z digest=sha256:ba6dffad24d5069eb08354b5ed75223bfda1777415d001e77fe3084e712466eb

Observation 37261bbf-4ae6-4b7e-a584-ed7e70e96505 · outbound

This paper cites SuperBPE: Space Travel for Language Models.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning SuperBPE: Space Travel for Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.770462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.770462Z digest=sha256:18297e9a4682f2cedb826690135cfbf05ca79d5d95eec81cdd318fc36627b537

Pith citing papers

Observation 744632db-d0d0-4464-9d5d-2d4e7ddb6978 · inbound

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers cites this paper.

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:25:13.091782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:25:13.091782Z digest=sha256:9480cd7bd8413948d0b6afd1d85bce41031077b04287418ca6c7d8c8efd63e86

Observation e40915e6-15aa-48d3-b672-1eb04ce95436 · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:10.690941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:10.690941Z digest=sha256:b44424dd9e0631b3636248c3c75ef2106e768ece32d34a883130978d2266a749

Observation 566d3c34-cc63-4f5f-b2ea-a9eae8368a52 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.166923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.166923Z digest=sha256:6f128d96a8f8ff96865d9dbc885384ed7209025c00cd56fff25dded5540999e5

Observation a3175c71-e39c-4f06-bddb-def2e6391fd0 · inbound

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE cites this paper.

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:38:54.418329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T00:38:53.782733Z digest=sha256:44ecaf1b36544a55d673c46d91975702997a3e762908a0074dfef73034d7fdd7