Pith. sign in

Paper Citation Record · LEDGER

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 2 inbound Pith citation observations for arXiv:2502.07057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07057 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:55:55.418419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.438760Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8435d93-5282-4c6e-976b-0d38f7f10bbc · outbound

This paper cites Tokenization Is More Than Compression.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenization Is More Than Compression

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.892514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.514018Z digest=sha256:6f73097598613afcd6349962994d06fcc4507a32f51b9626b0eb2dfee100aafd

Observation d374156d-3e5b-47e0-86a6-10cce3c7338e · outbound

This paper cites an unresolved cited work.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:58:58.881244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.518772Z digest=sha256:abe8ab384c00e6bf6b550718e53cab08390e8cc03e451d01f7f5721ff154e44a

Observation ddbab64a-6d8f-43eb-bb42-51dfb39d4943 · outbound

This paper cites How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.522901Z digest=sha256:bf4824830c80174ffb6dd09089de3f3d09c315ad9ab0aee2fd3dfdff7d7bbee4

Observation 4da0aa48-8e4f-490b-9630-431a09e54865 · outbound

This paper cites Critical tokenization and its properties.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Critical tokenization and its properties

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.857711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.526306Z digest=sha256:f48f12c70bf686d544fab27e43af74615e291b86a67e58ec4026c04f5fff6747

Observation 5f10b29d-b30a-42b0-bc2f-ae8cdf34b639 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.530907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.530907Z digest=sha256:3895b909567b860927bed8aa18ee476865dcc0fbefa7ec4893047118411f81bb

Observation 5926943e-bc6e-4ab8-afe4-06ef4863c5d4 · outbound

This paper cites Chemotactic motility-induced phase separation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Chemotactic motility-induced phase separation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.535300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.535300Z digest=sha256:1889341cc3fde9321ad00c0c8a677be8dce28e62653e9f6197b6ee84cb25faac

Observation 2f903b61-5838-4090-a979-bd945a6b9c11 · outbound

This paper cites Formalizing BPE Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Formalizing BPE Tokenization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.539958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.539958Z digest=sha256:5dcd62212f221ef34302152bbcd6eb369bea198f1fb26bca1ff27a4104c16ac3

Observation af1abc08-fbc1-4fbf-aa29-71647036cd41 · outbound

This paper cites The Technical User’s Introduction to LLM Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark The Technical User’s Introduction to LLM Tokenization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.846367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.544214Z digest=sha256:fd38c107f5b43b0ee9112ace96c766ded2ee5e076ef09c361e4e4bc802f56fea

Observation 3335b46e-21c7-45a8-9e96-c7d68c67edff · outbound

This paper cites A New Algorithm for Data Compression, 1994.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark A New Algorithm for Data Compression, 1994

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.834784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.548336Z digest=sha256:fdee4db379adb99a295cb9304cfd478937287d56c8b4a3e7a15170116efac5bb

Observation 4a6da132-4bf4-4a48-881b-c155a60e938c · outbound

This paper cites github.com/riotu-lab/aranizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark github.com/riotu-lab/aranizer, December 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.822901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.551338Z digest=sha256:9a0aff45ddd3221f91177e308c2520365706e3b7b60e829de2b666883c2ad787

Observation 06daf0fe-fe4f-4285-82de-016854153ab7 · outbound

This paper cites So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.810188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.554589Z digest=sha256:98671c22a88e63ae6f6e5b4c3cf809e9f22b9ad53ced949c0b0ad6d0b724c370

Observation 1e9a9452-5499-4bf5-905d-83e47c234d98 · outbound

This paper cites Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.793754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.558056Z digest=sha256:ecb13e89db8c0c65fd2767a2c78fb3d6283c1b2ade57461be3a61feed2b73baf

Observation fb2149e2-4a93-417c-994c-ca0684d2ef07 · outbound

This paper cites NbAiLab/tokenizer-benchmark, November 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark NbAiLab/tokenizer-benchmark, November 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.782283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.561881Z digest=sha256:c654029fec76cb02f448d167bedd9bfcc2a788c69c6f63a1f053d7dbf9124839

Observation 099c238d-30de-48db-8698-9fb6448049b4 · outbound

This paper cites Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.768447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.565542Z digest=sha256:b88e9ce991eabb1e1745450d87dc0739e03814f052562fd14552e62507935d3a

Observation 128c76ba-c58e-4c16-bb51-c43d74aab12c · outbound

This paper cites Analysis of Subword Tokenization Approaches for Turkish Language.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Analysis of Subword Tokenization Approaches for Turkish Language

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.754936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.568407Z digest=sha256:a02503661f781026145b32555bd127aa3b0a1888d2f5a2e7f62e525bf1bfa7bc

Observation 89aab7e3-7877-4b3d-9daf-13b67e6a345f · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark EuroLLM: Multilingual Language Models for Europe

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.571329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.571329Z digest=sha256:d1dc184b3a0af52cfa7fe287159a767c431eda3c66c3b4ae84490d4d8984b225

Observation 2a8a8885-e358-415a-b3dd-2f010c29951e · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.742478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.575667Z digest=sha256:05b5f81ec87c472dad2b59b28e0307755f544cafa12180c1e80010a821a4104c

Observation da76be89-48bc-459d-93d1-f78fcf1932b6 · outbound

This paper cites Not All Tokens Are What You Need for Pretraining.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Not All Tokens Are What You Need for Pretraining

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.730830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.583279Z digest=sha256:e8d27926a0671f84d7f97cfdc21704dc96bbf8d53a3afd8054c40bba4972c0f3

Observation 2cfa602c-f3d2-4591-929f-7b3963a6b6ef · outbound

This paper cites Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.586474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.586474Z digest=sha256:65f6efbe1bbb9f60cff37b624c19c27b486717d59b1eea6a091189e16242cf81

Observation a9daa4db-dbf0-47ca-b0c4-818c15cd2366 · outbound

This paper cites ITU Turkish NLP Web Service.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ITU Turkish NLP Web Service

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.718416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.590224Z digest=sha256:b7d03affa37f753a27d2cd91d8249ee2a768b8561458badad054a29dc9c9570c

Observation 617aa5f9-e6b1-4558-b0ff-741e6ce74e0b · outbound

This paper cites ahmetax/kalbur, October 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ahmetax/kalbur, October 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.706260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.593375Z digest=sha256:48eedf15693356134b8e8486f88806fe30c8bb2e0feb9e8733c2f16fd5d74b56

Observation 10db0fcf-e5b1-4566-a060-6c33370a1828 · outbound

This paper cites Ali Bayram.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Ali Bayram

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.694284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-08T13:58:58.597059Z digest=sha256:d206d0ba95134fae8d80239ebead219d6d3ca5d32e3eded72f057b5149a7e5d7

Observation 87669eec-12e7-4ce2-b555-855d4fbc350a · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.579433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.579433Z digest=sha256:035df9f5235effb3254f6fb01cb40370cbd1945aa7f44bfe42567a4c7a9468a1

Pith citing papers

Observation a092e4ed-44e1-46e1-b0ae-cb18ff71aaa6 · inbound

Tokenization Matters: Improving Zero-Shot NER for Indic Languages cites this paper.

Tokenization Matters: Improving Zero-Shot NER for Indic Languages Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:55:55.418419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:55:55.418419Z digest=sha256:364edb4bae07308539b4c19e3b7aab34b1a9d7a20f359f336ac3025a59b4c6b4

Observation 9c59362a-d240-4fc7-96b8-55705f8931c1 · inbound

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish cites this paper.

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.441232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T20:53:57.646472Z digest=sha256:4fe7d65d4d389f0f8ab44d6c49422d324703ddd6f4f4b3aee435efe715397377