Pith. sign in

Paper Citation Record · LEDGER

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

As of 17 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 4 inbound Pith citation observations for arXiv:2505.09738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.09738 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:30:03.795718Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:25:13.091782Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T00:38:54.307688Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e2aa0f6-8172-4036-88b6-6c6167d80146 · outbound

This paper cites ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.740361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.740361Z digest=sha256:e6e23433eaead539e3ec79250d00dbf18afcc43a483b310ecd3c03a3486718a0

Observation 22befeec-313f-4062-ac01-362ad341eb28 · outbound

This paper cites An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning An Empirical Study on Cross-lingual Vocabulary Adaptation for Efficient Language Model Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.748757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.748757Z digest=sha256:f27552f2bf9ea506d294d930150cf033dfc38ea55a3db0368c575bf24061f0a2

Observation 73b21c49-599b-4ed1-a22f-805c7da9fc19 · outbound

This paper cites Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.753278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.753278Z digest=sha256:15bd33e55e9489e4dd48c2d5eec17830b6165a0e2464c18269dc02d1f12372f5

Observation dd926726-e4b4-429c-bd75-a986b2989689 · outbound

This paper cites Airavata: Introducing Hindi Instruction-tuned LLM.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Airavata: Introducing Hindi Instruction-tuned LLM

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.758013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.758013Z digest=sha256:3059311c5aaa66206e878ca28aaee9c8b86727f6969cd802e39d563fd1850083

Observation e6088e76-f291-4698-8aa2-21c1e9409625 · outbound

This paper cites Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Trans-Tokenization and Cross-lingual Vocabulary Transfers: Language Adaptation of LLMs for Low-Resource NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.762381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.762381Z digest=sha256:ac72b4b1c73287bd513d7f76aabd2219d047b0d583bd3fbc23ab5f9b1605d94a

Observation 6239edb4-55f0-4240-ae0b-84260c0ed6d1 · outbound

This paper cites WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.774027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.774027Z digest=sha256:716cd1225148a76bcc2105b78df6058888b73ec25e1cb9f92336f72810c980a9

Observation 8ac63071-f400-43c9-9dec-3435350947f5 · outbound

This paper cites Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Pretraining Vision-Language Model for Difference Visual Question Answering in Longitudinal Chest X-rays

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.777814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.777814Z digest=sha256:9dceefebdd3a3919aa02fbdb2cc7f9ef05e4b302401b9b765a2f396d0e81d1aa

Observation be618465-4cbb-41db-b90e-2ecc74ea96af · outbound

This paper cites The ultraspherical rectangular collocation method and its convergence.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning The ultraspherical rectangular collocation method and its convergence

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:30:03.877668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:30:03.786577Z digest=sha256:0f6d720eef8604dc11ef834c95b4e061cedc7ea8be37ed6568545e29979428c4

Observation b0a10a52-e42c-40da-95eb-fcbe414bd2c3 · outbound

This paper cites Attention Is All You Need.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Attention Is All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.791522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.791522Z digest=sha256:58420cc4b2d89f587d2d826f82e1751cbe8755c7364b9d8ebf4ffd7c80240e35

Observation 9b6c678e-1e54-4153-b26a-4cde16206477 · outbound

This paper cites Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Autonomous Data Selection with Zero-shot Generative Classifiers for Mathematical Texts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.795718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.795718Z digest=sha256:e572ed010fe8f4f278bc5596134621249f44d0d508a995d936e864f02361b132

Observation 315fa27d-32fe-4ab9-88b4-804ca606584a · outbound

This paper cites doi: 10.18653/v1/ P16-1162.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning doi: 10.18653/v1/ P16-1162

Reference 2016

Resolution
malformed identifier
no resolver link, observed 2026-08-15T21:30:03.782214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.782214Z digest=sha256:089e2a7167f25b9e089b44b8540a36cea294e5b1cebf07ae82231b2b9ecdd874

Observation 67aba5da-2eeb-4886-8c77-8bd1c2247c84 · outbound

This paper cites an unresolved cited work.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:30:04.015946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:30:03.766499Z digest=sha256:d26a6feb2bf6886e176af96db2e917d2df915aa181d57b2ebd7d36728c0e3a7d

Observation 644ef3b4-a3a6-4d1c-9612-bfde45b66f07 · outbound

This paper cites Tamil-Llama: A New Tamil Language Model Based on Llama 2.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Tamil-Llama: A New Tamil Language Model Based on Llama 2

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.736366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.736366Z digest=sha256:5e4f9f259d4ccdf22243c6cc3815a6f3627e0fbfa8e1f37717bc9e8284973233

Observation 07b60db6-c8cc-41de-b76e-b130947d0468 · outbound

This paper cites Entanglement generation in capacitively coupled Transmon-cavity system.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning Entanglement generation in capacitively coupled Transmon-cavity system

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T21:30:04.003902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:30:03.731822Z digest=sha256:3e185997d4ca09e3f5d2413b2a3a404499bc308b6124494f2b0a307b549b5b18

Observation 37261bbf-4ae6-4b7e-a584-ed7e70e96505 · outbound

This paper cites SuperBPE: Space Travel for Language Models.

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning SuperBPE: Space Travel for Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T21:30:03.770462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:30:03.770462Z digest=sha256:560a6c2f4fbc2dd466b4c00d7695fbe79641245c8f0c1dcae35fba337b78f01b

Pith citing papers

Observation 744632db-d0d0-4464-9d5d-2d4e7ddb6978 · inbound

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers cites this paper.

One Tokenizer To Rule Them All: Emergent Language Plasticity via Multilingual Tokenizers Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T04:25:13.091782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:25:13.091782Z digest=sha256:b71eb8ce12e4790b0ce9f88dc36c0b39e23fac355c25582cee5be188a9c62a9b

Observation e40915e6-15aa-48d3-b672-1eb04ce95436 · inbound

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems cites this paper.

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T22:58:10.690941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:58:10.690941Z digest=sha256:3d85ccb9818b1b13b29e68a572c44bf13bba3750a787fa9ed7d14ef4833c8682

Observation 566d3c34-cc63-4f5f-b2ea-a9eae8368a52 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:17.166923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:17.166923Z digest=sha256:596f9453c7da0fc4e50a775eb81063b5c3c9263fa2381d42475e54fd6afbb7ae

Observation a3175c71-e39c-4f06-bddb-def2e6391fd0 · inbound

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE cites this paper.

Writing-System-Level Tokenizer Adaptation for Byte-Level BPE Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

Reference 2021

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T00:38:54.418329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T00:38:53.782733Z digest=sha256:09866b05242ef69cac2ba521f268a37f0ca3a06addc2162f7d397ff37fb8cff2