Pith. sign in

Paper Citation Record · LEDGER

NITP: Next Implicit Token Prediction for LLM Pre-training

As of 5 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.24956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24956 v3

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T18:45:28.635910Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f872ea8d-66a4-4c18-858f-a13568cdc3d2 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

NITP: Next Implicit Token Prediction for LLM Pre-training GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:845b5964b0337f9049a3c6b9fd863bf129ca68edf87184c0351269c583cb7787

Observation 1c3ab298-d7ac-4ac3-b52e-158f88d19a75 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:3fe862b1c20c848370d39df10b8222c36aec7add853f397ae4eab5c1779ba80a

Observation 284b326c-502b-4174-ab73-84a170a6ca51 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

NITP: Next Implicit Token Prediction for LLM Pre-training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:536a0472d4370e0918432127cc156e4215d123c012c5d094edc23c84e6caa483

Observation 45c18772-fa57-40d2-b270-10e9b8b39c34 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NITP: Next Implicit Token Prediction for LLM Pre-training Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:b3ccd7634fcd12bc8bd3c6aa69a826d5c62951dbfe10f126078a5b71ac7a73a0

Observation 45296ac4-6d8c-4cf3-a99d-4716d05b9dfa · outbound

This paper cites How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings.

NITP: Next Implicit Token Prediction for LLM Pre-training How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:24cd7eefa65dfaf7f2d9b26cdf55f0cfa6e2a2e085f957869381c8d80b91264c

Observation fe23d654-343f-4ae3-83cb-cb1bc98aa189 · outbound

This paper cites Representation Degeneration Problem in Training Natural Language Generation Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Representation Degeneration Problem in Training Natural Language Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:331807920db3f7822635b7f98e3726c6eaf0fff3e94afabbf91ee6d2f43f81d1

Observation bedc496f-98f8-4828-a1bb-346767839919 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

NITP: Next Implicit Token Prediction for LLM Pre-training The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:5a14e3dcb53f961fe1ba103c0d0fd60a7be4efad32efb7d04f5e2c8e87e9dea7

Observation fee8a5db-d6e9-4976-98eb-aa11357ca289 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

NITP: Next Implicit Token Prediction for LLM Pre-training Better & Faster Large Language Models via Multi-token Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:c2eca5a70b5d40e37260c03fff21706c22c69d8f52e3440a2a185bce5a8f8ecb

Observation fb9c5d06-ff17-4927-bb5f-b367f5d17433 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:49fae5c43b657dfd364f75f76f48fcef7ce4475412d5c0b1f67e21895a9e86c5

Observation 20f24f6c-ba20-4f29-8194-7b684153a2dd · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d277760b3ad42503e065c2d395a505aa46f3d5dd3f2ac62710eabf8f7a5075ab

Observation 8c5556cf-97e5-4241-9fc3-5ca019c6b71f · outbound

This paper cites Training Compute-Optimal Large Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Training Compute-Optimal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:33e9676cee3600cc887543842ad0a04508816a4ea04c7a93cd218f684103a047

Observation b98441cd-c429-4287-9f12-a99ff671c262 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Tinybert: Distilling bert for natural language understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:27be38b4a813f408a3f7ad8bdde03ac065ad0818b9d8d69ff364900d3a6fcbf8

Observation 3b779c04-d0ac-48bf-bc8f-10e2a7ecb00a · outbound

This paper cites Scaling Laws for Neural Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Scaling Laws for Neural Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:dd1d1085f3651062223827324364fd1794bfdd1c3e220a80b8c4bfcb1cc6e2c8

Observation 23deae2c-26ff-444a-8d49-9ba1b3fce52b · outbound

This paper cites When Choosing Plausible Alternatives, Clever Hans can be Clever.

NITP: Next Implicit Token Prediction for LLM Pre-training When Choosing Plausible Alternatives, Clever Hans can be Clever

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:e942ca3d6edaf5c163a9271184fea6c32fef013b9e91d107863e0d4b3350a2d1

Observation 3099c140-0127-490d-8ac2-4375c0ded92b · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

NITP: Next Implicit Token Prediction for LLM Pre-training Self-Distillation for Further Pre-training of Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:e54bbe1460c7ad09258317c966b7af051a8d93c7a4651b66f9922666b005ed7d

Observation 4e076dbe-51b2-468d-9bd7-c5cea22c1e0a · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:76608429f5949d3de0f793b72f23d0aa3e8f5f27dbff2cab8dd498d03a830baf

Observation e08aa440-a3fe-4438-93b4-fc33196a1850 · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

NITP: Next Implicit Token Prediction for LLM Pre-training Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:e48445b902b2cce985c3a1ffa603f8cf48bc1f7f982f16f20a99f633d9d67ed1

Observation bad0bc0a-3a26-41f8-b3fa-c9e66bb9b2ce · outbound

This paper cites Language Models are Few-Shot Learners.

NITP: Next Implicit Token Prediction for LLM Pre-training Language Models are Few-Shot Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:8890c07fdbb1c17eed275c459a9feb985e879579b0a496f4fc73f355e10d8e05

Observation 2099a6f5-6b24-4663-bab6-e0f1aa5c9ee5 · outbound

This paper cites Mteb: Massive text embedding benchmark.

NITP: Next Implicit Token Prediction for LLM Pre-training Mteb: Massive text embedding benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:81af735fdb344d816e39953db22d4e9da17821facae29aacde25fb6e25b86049

Observation 8ab55965-31d0-4884-9c49-0a38bbd37a0e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

NITP: Next Implicit Token Prediction for LLM Pre-training Representation Learning with Contrastive Predictive Coding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:303c7f784a5641169b8b3531609a75f05c66d104c50316460677e8e96d09bc5b

Observation 4e1aa5cd-9562-4bbb-ac68-0669e7eb8dde · outbound

This paper cites Future Lens: Anticipating Subsequent Tokens from a Single Hidden State.

NITP: Next Implicit Token Prediction for LLM Pre-training Future Lens: Anticipating Subsequent Tokens from a Single Hidden State

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:4489c36015d704a707045886b54d51eb1934089fc4ac932dded0a785dc25544b

Observation bd5c26b7-6ce6-4f8e-b151-da3a192d7300 · outbound

This paper cites FitNets: Hints for Thin Deep Nets.

NITP: Next Implicit Token Prediction for LLM Pre-training FitNets: Hints for Thin Deep Nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:ace04edee0028ffb836effe2f7f1a4439d05e77ee77bb5a1ad2cf4c5dd242e7b

Observation e912ef23-9fff-4880-a03c-c67b89d2f48e · outbound

This paper cites Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential.

NITP: Next Implicit Token Prediction for LLM Pre-training Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:85fde54ef686fcdd022e05af55474bfed0319cbd9f185347f95b4a4b14389e0e

Observation abf7d550-48e4-43ae-9c88-02639a54e9e0 · outbound

This paper cites GLU Variants Improve Transformer.

NITP: Next Implicit Token Prediction for LLM Pre-training GLU Variants Improve Transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:9d180578c272ad9548420b9c1b5889b38acccdd6f9ef980df59e2e310e97d66a

Observation 0a267b32-6956-43bc-af60-19367af57619 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

NITP: Next Implicit Token Prediction for LLM Pre-training Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:b80b5f4dd064e4b7d0c563c82de4e8cbdd9200a1f20eb98432eb4d97dcedbcc0

Observation 25fbb8c4-dd70-4ea6-a349-250c75b9e57b · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:2a6c928283cb2c67a511bc8b62a4408f0fdf28c8b10c8a5994241d0dabb154ad

Observation d7be940a-e20c-4407-add2-74746e405e6a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:15a3322554e82f29f453ededf10beacbdab08617074d71aab04c85ec86db0a2d

Observation 6867df25-a06b-4523-b272-5c9f14dc8d43 · outbound

This paper cites Contrastive distillation on intermediate representations for language model compression.

NITP: Next Implicit Token Prediction for LLM Pre-training Contrastive distillation on intermediate representations for language model compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:9bad149812762c25ecb69f5be2ab78de0a850bbdf6814981b7568e17dc8ea36e

Observation 80db826e-7060-4cb0-b06e-9e99428f908b · outbound

This paper cites LLM Pretraining with Continuous Concepts.

NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:26d25b22e53a560622fcfa0171b4312351bfd6ff07d20688f8eeda93d6311a6b

Observation 2bd389a7-40bc-461f-893c-3f7aa8870495 · outbound

This paper cites Com- monsenseqa: A question answering challenge targeting commonsense knowledge.

NITP: Next Implicit Token Prediction for LLM Pre-training Com- monsenseqa: A question answering challenge targeting commonsense knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d371c8ffbfaedc31f4b61d1876af07e314ee270a0052bdc5779b3821a85fb437

Observation d8b4d4d1-d6fc-4534-b532-4fab4cef1492 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

NITP: Next Implicit Token Prediction for LLM Pre-training Kimi K2: Open Agentic Intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d38e7eeece538cf151b7066ffe82840805abde8ac1f01f8be627aff85dadf59c

Observation a1f1d358-8aa6-49fa-8bdb-b54f577e9629 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

NITP: Next Implicit Token Prediction for LLM Pre-training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d4e473d47e17352c3dce6d6513b1499f1340183e7205255576a460aef4b2b50f

Observation 9badcd3a-81d1-49f7-9aa3-0f92107f0e51 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

NITP: Next Implicit Token Prediction for LLM Pre-training Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:6b7b24e3df9b82762b2e03c74896feaa04b4a41906c29b3060dae39a29d9e99a

Observation 1d4a81b6-8d2c-4bee-b12b-acec38b0ce69 · outbound

This paper cites Qwen3 Technical Report.

NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:140625b3b8abf60beaf5a5f03493b733df3df5cd69c5b810324f2298a6c8e9b8

Observation 5bf1e222-ee6d-4a1a-8ed2-0a47cad18eef · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

NITP: Next Implicit Token Prediction for LLM Pre-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:cf16dfaf7cce499965650554c05210be704fb1be83dfbaf3504fa216f17d642d

Observation b4e6b3dd-02c7-4bc0-9151-b766bc10cc32 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:61127ff3db3fd813beb69ee9942077f53d50a5f3cfe1c0e7d4c173a78ac38d5a

Observation 544d1b85-a7b1-43dd-a17b-48903d3d0421 · outbound

This paper cites Repre- sentation degeneration problem in prompt-based models for natural language understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Repre- sentation degeneration problem in prompt-based models for natural language understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:83ec95ade8888a2ebfa5b40b39f364da934ef070a6f570e8315c2bf29464e546

Observation 912ffb0b-1a9b-40b0-802e-9bbb1c9cf8e9 · outbound

This paper cites Agieval: A human- centric benchmark for evaluating foundation models.

NITP: Next Implicit Token Prediction for LLM Pre-training Agieval: A human- centric benchmark for evaluating foundation models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:7cf5922997da5e56b220797837faf93188b3cf4edaa43af40b025fc40bacac84

Observation 152b2086-2de2-4e05-925e-69cfe7c30b62 · outbound

This paper cites M., Fuadi, E.

NITP: Next Implicit Token Prediction for LLM Pre-training M., Fuadi, E

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:258041d91fd284b189f5e646ac55aa17b22d18c768af00e89f4d540090715d48

Observation 7cbb8c04-2c59-45aa-b79b-526ed788c2ff · outbound

This paper cites The learning rate and global batch size are scaled according to model size, while the context length is fixed to 8192 tokens for all experiments.

NITP: Next Implicit Token Prediction for LLM Pre-training The learning rate and global batch size are scaled according to model size, while the context length is fixed to 8192 tokens for all experiments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:757b0071c8b8a1331e727c249e5b20635c7fea5479b315e7545ca3d192fcf6c7

Observation 77a4a9a2-93e2-448e-b89f-02ed21f891e3 · outbound

This paper cites 2.006 (PPL 7.43 vs.

NITP: Next Implicit Token Prediction for LLM Pre-training 2.006 (PPL 7.43 vs

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:8374672cdedc69b25d9f81c6e6c962777f78ccb3ed1d29dacbbdedf263663b0d

Pith citing papers

No inbound Pith citation observations are available.