Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2502.07057.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:58:58.597059Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-26T20:53:57.646472Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:49:19.438760Z
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e8435d93-5282-4c6e-976b-0d38f7f10bbc · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenization Is More Than Compression
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d374156d-3e5b-47e0-86a6-10cce3c7338e · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ddbab64a-6d8f-43eb-bb42-51dfb39d4943 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in japanese
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4da0aa48-8e4f-490b-9630-431a09e54865 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Critical tokenization and its properties
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f10b29d-b30a-42b0-bc2f-ae8cdf34b639 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5926943e-bc6e-4ab8-afe4-06ef4863c5d4 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Chemotactic motility-induced phase separation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f903b61-5838-4090-a979-bd945a6b9c11 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Formalizing BPE Tokenization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af1abc08-fbc1-4fbf-aa29-71647036cd41 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark The Technical User’s Introduction to LLM Tokenization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3335b46e-21c7-45a8-9e96-c7d68c67edff · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark A New Algorithm for Data Compression, 1994
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a6da132-4bf4-4a48-881b-c155a60e938c · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark github.com/riotu-lab/aranizer, December 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06daf0fe-fe4f-4285-82de-016854153ab7 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark So many tokens, so little time: Introducing a faster, more flexible byte-pair tokenizer, December 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e9a9452-5499-4bf5-905d-83e47c234d98 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Arabic Tokenizers Leaderboard - a Hugging Face Space by MohamedRashad
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb2149e2-4a93-417c-994c-ca0684d2ef07 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark NbAiLab/tokenizer-benchmark, November 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 099c238d-30de-48db-8698-9fb6448049b4 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Tokenizing on scale.Preprocessing large text corpora on the lexical and sentence level
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 128c76ba-c58e-4c16-bb51-c43d74aab12c · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Analysis of Subword Tokenization Approaches for Turkish Language
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 89aab7e3-7877-4b3d-9daf-13b67e6a345f · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark EuroLLM: Multilingual Language Models for Europe
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a8a8885-e358-415a-b3dd-2f010c29951e · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models, June
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da76be89-48bc-459d-93d1-f78fcf1932b6 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Not All Tokens Are What You Need for Pretraining
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2cfa602c-f3d2-4591-929f-7b3963a6b6ef · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9daa4db-dbf0-47ca-b0c4-818c15cd2366 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ITU Turkish NLP Web Service
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 617aa5f9-e6b1-4558-b0ff-741e6ce74e0b · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark ahmetax/kalbur, October 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 10db0fcf-e5b1-4566-a060-6c33370a1828 · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark Ali Bayram
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87669eec-12e7-4ce2-b555-855d4fbc350a · outbound
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c59362a-d240-4fc7-96b8-55705f8931c1 · inbound
Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.