Pith. sign in

Paper Citation Record · LEDGER

Language Model Tokenizers Introduce Unfairness Between Languages

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2305.15425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.15425 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:32:25.383880Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

29
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5493db6b-1d07-41e4-a038-85a36e15f45b · inbound

Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models cites this paper.

Multilingual Question Answering in Low-Resource Settings: A Dzongkha-English Benchmark for Foundation Models Language Model Tokenizers Introduce Unfairness Between Languages

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:25.383880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:25.383880Z digest=sha256:1e43e9655a15a2ca63d06c2358ab2beaf2ebf0b7b47a032abd55036268f447ea

Observation 1121a55b-fb62-46b6-b0d0-c893765cad1f · inbound

Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts cites this paper.

Less, but Better: Efficient Multilingual Expansion for LLMs via Layer-wise Mixture-of-Experts Language Model Tokenizers Introduce Unfairness Between Languages

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:00.875465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:12:00.875465Z digest=sha256:69c66e66d5553e5bb0b1b21004fc8a6fce4474a229a2389e40e103086149fc25

Observation 2108986b-cfda-40b9-b2b3-3d827f213f76 · inbound

Causal Estimation of Tokenisation Bias cites this paper.

Causal Estimation of Tokenisation Bias Language Model Tokenizers Introduce Unfairness Between Languages

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:43.589887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:15:43.589887Z digest=sha256:d512cc5346d66f6165c65d906a173b0668a819d62fe3b36a294765e360e47126

Observation a4448f31-a4c7-4396-90ac-f0d963686242 · inbound

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala cites this paper.

Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala Language Model Tokenizers Introduce Unfairness Between Languages

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:37:53.013482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:36:27.421985Z digest=sha256:a3b5d8707c4639e234b30dfee47c542ebcabb73f53a46decc93a4eec7f5811f1

Observation 487f9c6b-8106-4cad-9807-69ef24b82406 · inbound

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization cites this paper.

ONTO: A Token-Efficient Columnar Notation for LLM Input Optimization Language Model Tokenizers Introduce Unfairness Between Languages

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.781850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:21:56.826488Z digest=sha256:db8492b16c8baef0dd29524f1d3cc52d5713d49ed251e2fe579328aef7b78a65

Observation f4f19e13-e5cb-451f-860c-91646d30f46c · inbound

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study cites this paper.

Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study Language Model Tokenizers Introduce Unfairness Between Languages

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:09:02.558978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T21:08:44.929762Z digest=sha256:cdbeac142ff2338dbb8aeeb4dbde276b3cbf6d536959d3379748154b2697684f

Observation 54ff702d-6fd9-47d9-af6c-f107095d4e8e · inbound

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations cites this paper.

LLMs in Qualitative Research: Opportunities, Limitations, and Practical Considerations Language Model Tokenizers Introduce Unfairness Between Languages

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T16:03:33.014552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:59:56.231210Z digest=sha256:fff85322d29f882fb0cd16cc06d43258852f420b4d126cf570ee20d12a60299b

Observation fb7e9804-4133-4a84-a9e0-8ffa26f9554d · inbound

Tokenization with Split Trees cites this paper.

Tokenization with Split Trees Language Model Tokenizers Introduce Unfairness Between Languages

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:51:08.838187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T05:48:29.391178Z digest=sha256:391ec0e1ea0bfef9bfb6a78da009758229ede8dfdf0f994c35d3073796e23dbd

Observation 899c837d-d4cd-46af-94a5-8b9f213d26de · inbound

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty cites this paper.

The Tokenizer Tax Across 25 European Languages: Domain Invariance, Cross-Lingual Few-Shot Effects, and the Ukrainian Penalty Language Model Tokenizers Introduce Unfairness Between Languages

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:14:40.909415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:08:33.172423Z digest=sha256:275cc9b8c38243ee25a3a1ab5d65e2c3b973dcb4f07c785464f839d6419ad8cc

Observation b2321e79-237a-41f1-a0db-c5d3bd003ee3 · inbound

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base cites this paper.

BrahmicTokenizer-131K: An Indic-Capable Drop-In Replacement for o200k_base Language Model Tokenizers Introduce Unfairness Between Languages

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:13:14.768133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:12:16.235649Z digest=sha256:a1064dbe50d05c08fde4a530002adea732c4166e52877260e20a6ba1b251734d

Observation 69bdf272-e63f-4d5d-9a5d-062dab6f10de · inbound

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators cites this paper.

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators Language Model Tokenizers Introduce Unfairness Between Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:01:38.447490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:01:38.447490Z digest=sha256:4bd3dc05d1b9542f42e30b2793cc2ee6eeac37849bd20345ee991be3c90b3949

Observation 8afba8f7-16bf-4d88-83a8-5fab1838c950 · inbound

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis cites this paper.

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis Language Model Tokenizers Introduce Unfairness Between Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T23:49:11.023254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:49:11.023254Z digest=sha256:ec46c2bc7235f0c67974d90de1fbc40f48da56bdd0564aea875d179783035c0f