Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:47:59.758998Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2506.15138.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:47:59.758998Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 85669550-fee2-48d9-8eb3-2cd6c9637e2c · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2789bf8d-326b-439d-90f7-2be9a4f7d70f · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bae2c230-d894-4273-8add-201038cd7414 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e77dd50-a44c-48a7-b785-bd61375af239 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Adaptive BPE Tokenization for Enhanced Vocabulary Adaptation in Finetuning Pretrained Language Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cffd4828-c8e9-46fe-b940-8aa8ad98e616 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models An Expanded Massive Multilingual Dataset for High-Performance Language Technologies (HPLT)
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 733101f7-8814-4f71-8912-486b9cbbf11c · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db4457e3-6252-4afe-9da5-762f96a9ae02 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Getting the most out of your tokenizer for pre-training and domain adaptation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 341c825e-1042-437d-a958-29bc7bbea7c7 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a21da18-f050-4a52-be61-65550e9a292c · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models How do different tokenizers perform on downstream tasks in scriptio continua languages?: A case study in Japanese
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f92c986-c3c4-4fba-9b24-99655f8dd1ba · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e8ff8d0-00ac-4026-a0e4-52813f9448d7 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f6d8bf5-a482-459c-8c63-882cf13e938a · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models The Llama 3 Herd of Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4cec494-d386-4138-b3c1-2a5264af483b · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78465ae-67c3-4c52-ab8c-63d03f45f7c2 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c3cf55d2-6958-439d-b7c5-068a4e2dd097 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73793c27-5a6b-456f-a8ca-52424b03b142 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b62926b5-3e24-4dae-9b53-455cb89f8c62 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models What Changes Can Large-scale Language Models Bring? Intensive Study on HyperCLOVA: Billions-scale Korean Generative Pretrained Transformers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80bdba0-f3f6-4d44-b8a4-2ddea4b65272 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Subword Regularization: Improving Neural Network Translation Models with Multiple Subword Candidates
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b715a8-d174-4b1e-b7c2-defd639c9e5a · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f2780a0-6ce0-46ae-96f3-b5a5fc9d8b95 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 270361d5-5b98-4203-8fac-23144bed1b1a · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2dc0644f-1ccb-4f32-8429-fc5979070045 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98c41c12-dbd7-420c-9dae-edc9b3992592 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models DeepSeek-V3 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62e473fe-81fb-4462-af56-a8d1079d6426 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b87be74-0718-49c4-857b-bfd937bc2e48 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Decoupled Weight Decay Regularization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcd4ba1f-f441-4d7b-8a81-415fab548fc0 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Meta Platforms
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 30f439af-c322-42e0-8f82-ad13eecbb48f · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Between words and characters: A Brief History of Open-Vocabulary Modeling and Tokenization in NLP
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66be9740-471d-4b4d-8591-69fb407bb3b4 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models An Empirical Study of Tokenization Strategies for Various Korean NLP Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c7ce041-7597-47e7-9c25-1d27e7db5d9d · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12210d51-8fdc-475c-9b85-d63debcaca45 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models BPE-Dropout: Simple and Effective Subword Regularization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3d8d42-3a1e-4430-a935-f97cf2bc891e · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Zero Bubble Pipeline Parallelism
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b095c7b8-2a11-43a3-8012-bdd2f98125ce · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede2af3e-0a26-4c52-aa60-30649f6edb19 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0920566-f3ef-480d-874d-79fc57e6e13e · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48cd711-45f0-4228-95ad-5fb375a59087 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a844923e-b48f-4ee3-af77-939855f6b295 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Tokenization Is More Than Compression
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 663b83ba-8f9e-448f-974b-c2b492946fe7 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b6daecf-4ce1-44eb-874a-b68605080614 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Neural Machine Translation of Rare Words with Subword Units
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28a55d12-84f1-42a4-a04f-6251fa34c5d1 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06282b0-b9af-4278-8eae-5bbe90737881 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models GraphBPE: Molecular Graphs Meet Byte-Pair Encoding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede9fb37-7323-40da-ad93-437a2ae4f719 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c2ff89f-1880-4dc8-9817-8bdced36407b · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b7e52288-1eb2-47fa-8dd0-165a107a50c5 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c729b015-b5da-4362-854c-f21b28201505 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ca33ab7b-9023-43de-9866-0836d7d2750a · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62662ad5-5887-4af7-b6a9-c466c1c5c735 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Incorporating Context into Subword Vocabularies
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61c4df7f-8b48-41f6-8eb7-1eee03035ef1 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97e3abbc-fa98-4126-a164-b72201ff0742 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4c745d4c-f867-4088-8ba9-09f34a09e274 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models URL: " 'urlintro :=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a331d249-8366-463f-a9a3-1fa6767d7d33 · outbound
Thunder-Tok: Minimizing Tokens per Word in Tokenizing Korean Texts for Generative Language Models write newline
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.