Pith. sign in

Paper Citation Record · LEDGER

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

As of 9 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.07057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07057 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T20:53:57.646472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.438760Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8435d93-5282-4c6e-976b-0d38f7f10bbc · outbound

This paper cites Tokenization Is More Than Compression.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenization Is More Than Compression

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.892514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.514018Z digest=sha256:4a6a8866e7a837f63e5a0db3bdf08936b2428c7f8be6769230f0ef32cad901a6

Observation d374156d-3e5b-47e0-86a6-10cce3c7338e · outbound

This paper cites an unresolved cited work.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:58:58.881244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.518772Z digest=sha256:f8c19b9f1a4bfbaaf316fff1e5b178a1e4ddb9e4e0863cb8b646318c95214a95

Observation ddbab64a-6d8f-43eb-bb42-51dfb39d4943 · outbound

This paper cites How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.522901Z digest=sha256:d3d6b2d8c3f8b8f44eb3a0b71968f9790ba85f447a25f7f20121245c60e80c07

Observation 4da0aa48-8e4f-490b-9630-431a09e54865 · outbound

This paper cites Critical tokenization and its properties.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Critical tokenization and its properties

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.857711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.526306Z digest=sha256:1d6191a2923efcf137c98bdb0587c501af9fa1fd5d260ba52a4c348368538cfa

Observation 5f10b29d-b30a-42b0-bc2f-ae8cdf34b639 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.530907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.530907Z digest=sha256:b93bc3119040a7613d410ec973b1ca28f46ba3e85f5ea3123749c41bb60b1dbd

Observation 5926943e-bc6e-4ab8-afe4-06ef4863c5d4 · outbound

This paper cites Chemotactic motility-induced phase separation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Chemotactic motility-induced phase separation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.535300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.535300Z digest=sha256:5add5ffc8f24006fc78cb5a1aefb681b5651be312ff2bdf295d30df8fc75f2ec

Observation 2f903b61-5838-4090-a979-bd945a6b9c11 · outbound

This paper cites Formalizing BPE Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Formalizing BPE Tokenization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.539958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.539958Z digest=sha256:7e04f438f02a5cd4f66777e087db5496c7ec1e46b301f26e37cef0f44deb7137

Observation af1abc08-fbc1-4fbf-aa29-71647036cd41 · outbound

This paper cites The Technical User’s Introduction to LLM Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark The Technical User’s Introduction to LLM Tokenization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.846367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.544214Z digest=sha256:481ed8767aea6af89da2486762f2eab3bbce50e2131d1f5cae384b862c5dc374

Observation 3335b46e-21c7-45a8-9e96-c7d68c67edff · outbound

This paper cites A New Algorithm for Data Compression, 1994.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark A New Algorithm for Data Compression, 1994

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.834784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.548336Z digest=sha256:5273d2af704cad33e6a664f8a76e901c34c4e5ccb174338a1e3d25080f0295e1

Observation 4a6da132-4bf4-4a48-881b-c155a60e938c · outbound

This paper cites github.com/riotu-lab/aranizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark github.com/riotu-lab/aranizer, December 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.822901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.551338Z digest=sha256:961026c050988306309cfaa802bbdc77403530cc3cc9bd043af7096cfaf963d3

Observation 06daf0fe-fe4f-4285-82de-016854153ab7 · outbound

This paper cites So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.810188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.554589Z digest=sha256:dc5b45726fd995f33071535012f10d57c63e2674f8bc717c498cd6f19d7d6ca0

Observation 1e9a9452-5499-4bf5-905d-83e47c234d98 · outbound

This paper cites Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.793754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.558056Z digest=sha256:78b7b459ef031fe615c43053356d58385c9530706aa695a9c30d616457a5afd7

Observation fb2149e2-4a93-417c-994c-ca0684d2ef07 · outbound

This paper cites NbAiLab/tokenizer-benchmark, November 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark NbAiLab/tokenizer-benchmark, November 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.782283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.561881Z digest=sha256:4ddf039ddc2800d7288de457300e427f820720baa0901804acb1316f6f327dbd

Observation 099c238d-30de-48db-8698-9fb6448049b4 · outbound

This paper cites Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.768447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.565542Z digest=sha256:71885d54b547cf231854ac1a02884a36960c35e32528e2d7a024e55c3a4f5c9a

Observation 128c76ba-c58e-4c16-bb51-c43d74aab12c · outbound

This paper cites Analysis of Subword Tokenization Approaches for Turkish Language.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Analysis of Subword Tokenization Approaches for Turkish Language

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.754936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.568407Z digest=sha256:7b0d0b8dcc156e0485e77e5d61a2f3625555b5581d3e7db1207163057c748500

Observation 89aab7e3-7877-4b3d-9daf-13b67e6a345f · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark EuroLLM: Multilingual Language Models for Europe

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.571329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.571329Z digest=sha256:64e36a17dc1461f400bcea1a0d5db7d8df66c440326d388e632ba7e83944b6d6

Observation 2a8a8885-e358-415a-b3dd-2f010c29951e · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.742478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.575667Z digest=sha256:603dc05e4366690bae28a29506880a695c6b0d435cba5148470c01e080198072

Observation da76be89-48bc-459d-93d1-f78fcf1932b6 · outbound

This paper cites Not All Tokens Are What You Need for Pretraining.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Not All Tokens Are What You Need for Pretraining

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.730830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.583279Z digest=sha256:0f78632e30bcc70a68fe149617c3a6c05eb4ecc1fa94f0311bd0afc7358b0d21

Observation 2cfa602c-f3d2-4591-929f-7b3963a6b6ef · outbound

This paper cites Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.586474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.586474Z digest=sha256:6a273a577ab18df314fc4f3c7f135e605070f3b5503eb255f65d6e8992141f28

Observation a9daa4db-dbf0-47ca-b0c4-818c15cd2366 · outbound

This paper cites ITU Turkish NLP Web Service.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ITU Turkish NLP Web Service

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.718416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.590224Z digest=sha256:d16f6554979fd4d0eb0161239ed88b64efeaa80774ec90963235800aa1b509ca

Observation 617aa5f9-e6b1-4558-b0ff-741e6ce74e0b · outbound

This paper cites ahmetax/kalbur, October 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ahmetax/kalbur, October 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.706260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.593375Z digest=sha256:28ec0384b0579c9f9bd85decd76e0505550b96efab952369693b8f2d1122f699

Observation 10db0fcf-e5b1-4566-a060-6c33370a1828 · outbound

This paper cites Ali Bayram.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Ali Bayram

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.694284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.597059Z digest=sha256:46aa8aae94f98459f5772dee1b7cf5c343dcb39f887b47ff995aaa4767084367

Observation 87669eec-12e7-4ce2-b555-855d4fbc350a · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.579433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.579433Z digest=sha256:b65c1bc755268de62ac3f1021f6e8f477e4d741903246a699d227f5e4c69415f

Pith citing papers

Observation 9c59362a-d240-4fc7-96b8-55705f8931c1 · inbound

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish cites this paper.

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.441232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T20:53:57.646472Z digest=sha256:dadca119b2bfb1b80c2169c7340ff886379651ea086a00eb52cea91285b8470c