Pith. sign in

Paper Citation Record · LEDGER

Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2112.10508.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2112.10508 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:43.616119Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

106
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a1c99ac-44f9-4d62-887c-8c9e45f07472 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:11.075762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:7dbc9af1fde3919aafc38c46a078347868e3c4975155a5f7fc43f8bebc21abe8

Observation f105b10b-2fd9-4768-bee4-7a87a9141599 · inbound

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model cites this paper.

BLOOM: A 176B-Parameter Open-Access Multilingual Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 278

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:11.646731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T00:51:10.919818Z digest=sha256:988d738bad7f37b3598d7e2649c4635e540b94969a3313059527809cd2ef471e

Observation c204e047-8388-4312-9431-0d41cc1ddd0b · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:19:46.729797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:30577ffa52c330ead007feabed3e70c9ab9b5b3ace7f04560b78e3799cc36746

Observation d570476c-b086-43ce-89e5-2fdbafa6d6fa · inbound

A Comprehensive Overview of Large Language Models cites this paper.

A Comprehensive Overview of Large Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.607604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T20:28:38.900026Z digest=sha256:9c6cad0e3bf32809f15f9a614117cd1139bbf29fda166a468b7f8a48b5cbb9e6

Observation 4bd88b63-8ba7-4bb3-a279-175228a7692a · inbound

Jamba: A Hybrid Transformer-Mamba Language Model cites this paper.

Jamba: A Hybrid Transformer-Mamba Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:11:27.200041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T14:11:27.156350Z digest=sha256:d0db0cd3f9d8c6bbb069a31efb63d3d8ac16ad0eaceb631e25550d8f43ef5184

Observation 4f013cee-8e99-4990-b05f-77b42f49ad0b · inbound

Comparative analysis of subword tokenization approaches for Indian languages cites this paper.

Comparative analysis of subword tokenization approaches for Indian languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:43.616119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:43.616119Z digest=sha256:bcf0540d9cb7c5c26b8cf74667d61b40e3c1c21cf2b444ed5cf569b33bef5811

Observation 1f1860a8-eae6-4cc5-b178-70030171147b · inbound

Beyond Text Compression: Evaluating Tokenizers Across Scales cites this paper.

Beyond Text Compression: Evaluating Tokenizers Across Scales Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:18:11.036474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:18:11.036474Z digest=sha256:2ccb3a9a369a6d5d03636c99d1dae3ad852af4b2bea87ed2dccfe0f830fcafd0

Observation bde9fb3c-eb54-42c1-b85b-04165b01106a · inbound

AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data cites this paper.

AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:38.211148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:38.211148Z digest=sha256:a6a4e57082edeb89da35b937c06859d18fc9956dc7d9b0863a49a07207bf3003

Observation f176c6cb-890a-4371-b7b6-e53edebd78bf · inbound

Bit-level BPE: Below the byte boundary cites this paper.

Bit-level BPE: Below the byte boundary Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:47.776405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:47.776405Z digest=sha256:ffb9dca01c9d4681ed302e5c8260e68a46b4d2afa428d82e9027357706a9709c

Observation c4f9ba5f-e3f0-4321-a224-fd5e3daa375c · inbound

Is There a Case for Conversation Optimized Tokenizers in Large Language Models? cites this paper.

Is There a Case for Conversation Optimized Tokenizers in Large Language Models? Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:19:15.192201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:19:15.192201Z digest=sha256:05df39d7500461aef6c6ca024d9200cf7cdc093b2e8c5e35213a08e4c332e42c

Observation 30ff154b-ccfc-425f-815b-7b6da12380b0 · inbound

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language cites this paper.

FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:40.446794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:40.446794Z digest=sha256:ac80f34d71da18dd723127391d8f8db99e17145684ef7490a7291232a02aaaae

Observation a3c0260e-6aa5-4a87-8872-85a002c12540 · inbound

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models cites this paper.

Train It and Forget It: Merge Lists are Unnecessary for BPE Inference in Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:43:23.055609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:43:23.055609Z digest=sha256:20334c62e71cbab4137d1fa02d5385f47c31f249ae3d2666b06e5b8b08cf11f3

Observation 23e77333-f6c1-4321-9e02-8bce1b0b9835 · inbound

FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification cites this paper.

FlowletFormer: Network Behavioral Semantic Aware Pre-training Model for Traffic Classification Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:26:02.073354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:26:02.073354Z digest=sha256:bb39e510f6721e0539afffdd7e687e13aebc00cb852ab9b74b2f64caf295cfe9

Observation c7abf7d9-1473-425e-b7cf-f7e24b68acfd · inbound

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model cites this paper.

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T13:26:15.129449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:26:15.129449Z digest=sha256:10be85feeb033f9c00f98819092dd8b64880f0c35e1b633e74d29180fdd2c3f3

Observation abb38b1c-7949-4d7b-94ac-cfc58ebfb653 · inbound

Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning cites this paper.

Learning Mechanism Underlying NLP Pre-Training and Fine-Tuning Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T11:01:20.305515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:01:20.305515Z digest=sha256:c5299d66758b105e0f7346a26a4928285f2a0ce23d2a8b2b6cb821fcc3a71340

Observation aeb96fa4-aa26-40e8-9ff8-adabf20151cb · inbound

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali cites this paper.

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:49:36.205280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:48:32.432546Z digest=sha256:4a6b9f35874f0d72fc2539a6c78c25bcfd80df5969f9ed603f9b7e6bc6e062a3

Observation a2515f8e-e534-46a0-8c5e-42044d551339 · inbound

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources cites this paper.

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:01:02.017768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:22:34.046014Z digest=sha256:2a9080ef7d3823de989b7d10ddb0a229715c4fd8b7773e95f8a0ecfe47fd1636

Observation b47d1164-f5aa-4c99-91d0-46eeade1d351 · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:38:21.835952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:571c946eee8839c142b567940f1856354661127f7d9afa0b15bad2e5cb7b2ea4

Observation 93d585ba-5dcb-4b0c-8628-a416be556035 · inbound

Translating Signals to Languages for sEMG-Based Activity Recognition cites this paper.

Translating Signals to Languages for sEMG-Based Activity Recognition Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:44:42.857609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T07:41:26.887209Z digest=sha256:14ffdfc230a72ee2d35c4e55cf6d28645ca6e5aeb34eec2b7441cb3a86b12a84

Observation 533da98c-0bd8-4f52-a6a3-7915e716954c · inbound

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models cites this paper.

Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.399554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:50:25.019889Z digest=sha256:0f3f4495479b02bb27018fed150c762d13ad21329d98131c8a2f5dd32855b32d

Observation 85e7a8d4-9ca4-40a2-9acf-a8dfb0cd6e7d · inbound

The price of incrementality in k-center clustering cites this paper.

The price of incrementality in k-center clustering Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:28.310361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:41:41.502995Z digest=sha256:3680046fc2e9f3f143c70c511248b78436a4c14d39562034c1ab02174e036183

Observation cab02254-afc5-4f27-95c4-0ee3846bfd24 · inbound

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet cites this paper.

Phonemes to the Rescue: Multilingual Tokenization Based on International Phonetic Alphabet Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:39:34.660865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T16:53:25.556222Z digest=sha256:b60806902e4d4e9f1794776a3e108b50deab11f2e18767d67a21c0a227952361

Observation 60b98f86-83b7-4010-af89-a83726b84097 · inbound

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment cites this paper.

MinGram: A Minimalist Unigram Tokenizer with High Compression and Competitive Morphological Alignment Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.297220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T04:59:05.057552Z digest=sha256:976363b64e526037d9f3500d118cf918fea14045f6efc07896a70fedbb6c41c5