Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T18:45:28.635910Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.24956.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-14T18:45:28.635910Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f872ea8d-66a4-4c18-858f-a13568cdc3d2 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c3ab298-d7ac-4ac3-b52e-158f88d19a75 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284b326c-502b-4174-ab73-84a170a6ca51 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c18772-fa57-40d2-b270-10e9b8b39c34 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Training Verifiers to Solve Math Word Problems
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45296ac4-6d8c-4cf3-a99d-4716d05b9dfa · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe23d654-343f-4ae3-83cb-cb1bc98aa189 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Representation Degeneration Problem in Training Natural Language Generation Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bedc496f-98f8-4828-a1bb-346767839919 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee8a5db-d6e9-4976-98eb-aa11357ca289 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Better & Faster Large Language Models via Multi-token Prediction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb9c5d06-ff17-4927-bb5f-b367f5d17433 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f24f6c-ba20-4f29-8194-7b684153a2dd · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Measuring Massive Multitask Language Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5556cf-97e5-4241-9fc3-5ca019c6b71f · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Training Compute-Optimal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b98441cd-c429-4287-9f12-a99ff671c262 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Tinybert: Distilling bert for natural language understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b779c04-d0ac-48bf-bc8f-10e2a7ecb00a · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Scaling Laws for Neural Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23deae2c-26ff-444a-8d49-9ba1b3fce52b · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training When Choosing Plausible Alternatives, Clever Hans can be Clever
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3099c140-0127-490d-8ac2-4375c0ded92b · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Self-Distillation for Further Pre-training of Transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e076dbe-51b2-468d-9bd7-c5cea22c1e0a · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e08aa440-a3fe-4438-93b4-fc33196a1850 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bad0bc0a-3a26-41f8-b3fa-c9e66bb9b2ce · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Language Models are Few-Shot Learners
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2099a6f5-6b24-4663-bab6-e0f1aa5c9ee5 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Mteb: Massive text embedding benchmark
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab55965-31d0-4884-9c49-0a38bbd37a0e · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Representation Learning with Contrastive Predictive Coding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1aa5cd-9562-4bbb-ac68-0669e7eb8dde · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd5c26b7-6ce6-4f8e-b151-da3a192d7300 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training FitNets: Hints for Thin Deep Nets
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e912ef23-9fff-4880-a03c-c67b89d2f48e · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abf7d550-48e4-43ae-9c88-02639a54e9e0 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training GLU Variants Improve Transformer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a267b32-6956-43bc-af60-19367af57619 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25fbb8c4-dd70-4ea6-a349-250c75b9e57b · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Layer by Layer: Uncovering Hidden Representations in Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7be940a-e20c-4407-add2-74746e405e6a · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6867df25-a06b-4523-b272-5c9f14dc8d43 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Contrastive distillation on intermediate representations for language model compression
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80db826e-7060-4cb0-b06e-9e99428f908b · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd389a7-40bc-461f-893c-3f7aa8870495 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Com- monsenseqa: A question answering challenge targeting commonsense knowledge
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b4d4d1-d6fc-4534-b532-4fab4cef1492 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Kimi K2: Open Agentic Intelligence
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1f1d358-8aa6-49fa-8bdb-b54f577e9629 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9badcd3a-81d1-49f7-9aa3-0f92107f0e51 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4a81b6-8d2c-4bee-b12b-acec38b0ce69 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Technical Report
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bf1e222-ee6d-4a1a-8ed2-0a47cad18eef · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4e6b3dd-02c7-4bc0-9151-b766bc10cc32 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 544d1b85-a7b1-43dd-a17b-48903d3d0421 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Repre- sentation degeneration problem in prompt-based models for natural language understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 912ffb0b-1a9b-40b0-802e-9bbb1c9cf8e9 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training Agieval: A human- centric benchmark for evaluating foundation models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 152b2086-2de2-4e05-925e-69cfe7c30b62 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training M., Fuadi, E
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbb8c04-2c59-45aa-b79b-526ed788c2ff · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training The learning rate and global batch size are scaled according to model size, while the context length is fixed to 8192 tokens for all experiments
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77a4a9a2-93e2-448e-b89f-02ed21f891e3 · outbound
NITP: Next Implicit Token Prediction for LLM Pre-training 2.006 (PPL 7.43 vs
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.