Pith. sign in

Paper Citation Record · LEDGER

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.07057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07057 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-26T20:53:57.646472Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:19.438760Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8435d93-5282-4c6e-976b-0d38f7f10bbc · outbound

This paper cites Tokenization Is More Than Compression.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenization Is More Than Compression

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.892514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.514018Z digest=sha256:37d477237f7c83f562ac24e745ef02ec85e458a0d23799f2b3d5a39684c4a045

Observation d374156d-3e5b-47e0-86a6-10cce3c7338e · outbound

This paper cites an unresolved cited work.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:58:58.881244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.518772Z digest=sha256:6dc5c7a7fc26d7542d7cf10a307342b146994a2e7757e07452cb77c7d18b8455

Observation ddbab64a-6d8f-43eb-bb42-51dfb39d4943 · outbound

This paper cites How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.869782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.522901Z digest=sha256:ab19147386c129619b64d5af18021f7b151641537ce3479c9e2369b50fa897c1

Observation 4da0aa48-8e4f-490b-9630-431a09e54865 · outbound

This paper cites Critical tokenization and its properties.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Critical tokenization and its properties

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.857711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.526306Z digest=sha256:a75bff3054158beddca0d9e7970ac11ba6f7aee402cd413299f1eaa9a63d4eb1

Observation 5f10b29d-b30a-42b0-bc2f-ae8cdf34b639 · outbound

This paper cites SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.530907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.530907Z digest=sha256:e09247f8b212fc8c6e832d193221177920b59c1d8474fb1abd5b64fb6846ff4b

Observation 5926943e-bc6e-4ab8-afe4-06ef4863c5d4 · outbound

This paper cites Chemotactic motility-induced phase separation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Chemotactic motility-induced phase separation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.535300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.535300Z digest=sha256:87a3d7d286549c6a62cca7b98ff10761d9e33df6f20f3a5fa9350a7d0ab504b5

Observation 2f903b61-5838-4090-a979-bd945a6b9c11 · outbound

This paper cites Formalizing BPE Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Formalizing BPE Tokenization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.539958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.539958Z digest=sha256:94ce20c1b5128ffe0c3259610eae96c98b33d51b0b676051616b00db87519a41

Observation af1abc08-fbc1-4fbf-aa29-71647036cd41 · outbound

This paper cites The Technical User’s Introduction to LLM Tokenization.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark The Technical User’s Introduction to LLM Tokenization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.846367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.544214Z digest=sha256:bbffe881f265850a3ebdbade5e925bf891b29e9758e0a84cee7fc57c97a9ea7d

Observation 3335b46e-21c7-45a8-9e96-c7d68c67edff · outbound

This paper cites A New Algorithm for Data Compression, 1994.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark A New Algorithm for Data Compression, 1994

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.834784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.548336Z digest=sha256:ba97f0739de2a861b940254947a3de4baec0ed203692cb632c1565b6a60794c6

Observation 4a6da132-4bf4-4a48-881b-c155a60e938c · outbound

This paper cites github.com/riotu-lab/aranizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark github.com/riotu-lab/aranizer, December 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.822901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.551338Z digest=sha256:6a5c94ea2f12b39f7b07c3092c05c732841e952c0ab41ec2d3986fd9dd126e69

Observation 06daf0fe-fe4f-4285-82de-016854153ab7 · outbound

This paper cites So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.810188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.554589Z digest=sha256:30c0a713f02d01980154094e10215973139998c63c79074ea5986f1c5bc5dcf6

Observation 1e9a9452-5499-4bf5-905d-83e47c234d98 · outbound

This paper cites Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.793754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.558056Z digest=sha256:63174c952fcdbc2110fea741a62be7e216832709d13601e5c680fb17b0652c3f

Observation fb2149e2-4a93-417c-994c-ca0684d2ef07 · outbound

This paper cites NbAiLab/tokenizer-benchmark, November 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark NbAiLab/tokenizer-benchmark, November 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.782283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.561881Z digest=sha256:5bc35ddf1eac0b9af2ebac547a64df510aa479dde852422f7fada1921c5b43c0

Observation 099c238d-30de-48db-8698-9fb6448049b4 · outbound

This paper cites Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.768447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.565542Z digest=sha256:3f6e9655a926c6e304bb1ab582baf3043a43b1b12cf33de34085d3eb13987beb

Observation 128c76ba-c58e-4c16-bb51-c43d74aab12c · outbound

This paper cites Analysis of Subword Tokenization Approaches for Turkish Language.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Analysis of Subword Tokenization Approaches for Turkish Language

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.754936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.568407Z digest=sha256:a9400769c8e4043ea01785cef8687cef0996ea9197eee9c4856eddd035cc8504

Observation 89aab7e3-7877-4b3d-9daf-13b67e6a345f · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark EuroLLM: Multilingual Language Models for Europe

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.571329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.571329Z digest=sha256:f00ba2e719a24b4ec0e7e9b5e5a4d8e817cfa83b47520756d266446151e24c33

Observation 2a8a8885-e358-415a-b3dd-2f010c29951e · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.742478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.575667Z digest=sha256:30bb1cc0e59d1add58b9b077065ea1dcac4d7658a48825d54960e0f10a9fa3e6

Observation da76be89-48bc-459d-93d1-f78fcf1932b6 · outbound

This paper cites Not All Tokens Are What You Need for Pretraining.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Not All Tokens Are What You Need for Pretraining

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.730830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.583279Z digest=sha256:9af4667f24cd135e74a6a2ad1c460974717c471cabbeaed33c0ee96adc0d55a3

Observation 2cfa602c-f3d2-4591-929f-7b3963a6b6ef · outbound

This paper cites Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.586474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.586474Z digest=sha256:95c3e833c683459748f0aacf27130ad4b358bb6618e32007009f54756190466b

Observation a9daa4db-dbf0-47ca-b0c4-818c15cd2366 · outbound

This paper cites ITU Turkish NLP Web Service.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ITU Turkish NLP Web Service

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.718416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.590224Z digest=sha256:94f8294534bfdfaf99a1de80d0cad1a61cbe6c7de101216355b91f2488182e6b

Observation 617aa5f9-e6b1-4558-b0ff-741e6ce74e0b · outbound

This paper cites ahmetax/kalbur, October 2024.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ahmetax/kalbur, October 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.706260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.593375Z digest=sha256:eefdf0df5b2beb34606d7241f2d1f472889a27fa1b2b547f03d2c4ac2f7e2e45

Observation 10db0fcf-e5b1-4566-a060-6c33370a1828 · outbound

This paper cites Ali Bayram.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Ali Bayram

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:58:58.694284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T13:58:58.597059Z digest=sha256:0a3d0623383082a2b78e2002aa1ae921e30b9de0169ac8d912b1814599cffc1b

Observation 87669eec-12e7-4ce2-b555-855d4fbc350a · outbound

This paper cites How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models.

Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T13:58:58.579433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:58:58.579433Z digest=sha256:efc64661e891de4e592c9e3f6c48d1e40239093616d5db7c68cd5d2f0b9f71c4

Pith citing papers

Observation 9c59362a-d240-4fc7-96b8-55705f8931c1 · inbound

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish cites this paper.

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.441232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T20:53:57.646472Z digest=sha256:c5a27bb4a81221464ee3e56db93ab6112339531b0049822b8aad11fa668475ed