Pith. sign in

Paper Citation Record · LEDGER

LLM Pretraining with Continuous Concepts

As of 22 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 19 inbound Pith citation observations for arXiv:2502.08524.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08524 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:54:39.421494Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.771507Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 70ce4d20-5606-489e-8464-8873fe115a0a · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LLM Pretraining with Continuous Concepts Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.269793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.269793Z digest=sha256:e9d0871c8894cb5c002faa020f3a4cedb0675229b8c60eb385b74b2654676943

Observation 6995f843-6bb6-462c-93e1-aff7e85b49e4 · outbound

This paper cites DeepSeek-V3 Technical Report.

LLM Pretraining with Continuous Concepts DeepSeek-V3 Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.280782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.280782Z digest=sha256:01706eb31d46a28b3cc154d233a7a914c78ba54f5551c19dcb9654c697764d88

Observation 2c57982a-0ba5-455f-b6df-0fdb59de392e · outbound

This paper cites The Llama 3 Herd of Models.

LLM Pretraining with Continuous Concepts The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.291932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.291932Z digest=sha256:bdf4f84d64870f6dc8983c236155b9fdabe2cf36cd6d947b82217849fa76367b

Observation ca7a76e4-8a62-44d9-87d0-cce56ebf2776 · outbound

This paper cites AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn.

LLM Pretraining with Continuous Concepts AssistGPT: A General Multi-modal Assistant that can Plan, Execute, Inspect, and Learn

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.297648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.297648Z digest=sha256:f4f5c5468b27b0d14659cdd7c29722d69eb2d703d82283d1fa7dd987f8448c53

Observation eb8c4dec-efd5-4ed0-b1cf-fff501116883 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

LLM Pretraining with Continuous Concepts Scaling and evaluating sparse autoencoders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.302858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.302858Z digest=sha256:a06cc41b098ed2ca315fec9312f025ca6e326bb551f9d6e6b260b0ca602f9c62

Observation 654413b9-704c-431e-849e-de4cb7ab02e5 · outbound

This paper cites MiniPLM: Knowledge Distillation for Pre-Training Language Models.

LLM Pretraining with Continuous Concepts MiniPLM: Knowledge Distillation for Pre-Training Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.307725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.307725Z digest=sha256:661cd27d9aa5c11fc5baeb7d22a05e044bf6d62b8e8ab1b262fcec2cc719cbaf

Observation 329e6b1f-9aaa-4b11-bf5e-19f78a067d63 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

LLM Pretraining with Continuous Concepts Training Large Language Models to Reason in a Continuous Latent Space

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.313421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.313421Z digest=sha256:ec2f838031890ed1168d7a5e93f2236dc564bedfc468d2d3aac3737601ba633d

Observation 2f08ce55-b6f8-410c-98bf-cc948defd65b · outbound

This paper cites Distilling the Knowledge in a Neural Network.

LLM Pretraining with Continuous Concepts Distilling the Knowledge in a Neural Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.319043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.319043Z digest=sha256:d3d76ee7cc18686e2dbdd3ee71321a60af8e0f36803ef337e0f5db6a3e1f9aed

Observation a6096342-a395-41a8-9d9f-303c102c5783 · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

LLM Pretraining with Continuous Concepts Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.324272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.324272Z digest=sha256:1a8f66dc3c50ef891f99df9dc9fd815938f3168cef30e03c31af1c2441ff757e

Observation 4e812e43-1633-4be2-9ae8-9156c926c1da · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

LLM Pretraining with Continuous Concepts Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.330610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.330610Z digest=sha256:f037a846b0640859479748c9bad7505822998dd96ac28f2408de037eda0b3d32

Observation d2836ea7-38e3-486b-a59c-24988a3151f6 · outbound

This paper cites Decoupled Weight Decay Regularization.

LLM Pretraining with Continuous Concepts Decoupled Weight Decay Regularization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.336491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.336491Z digest=sha256:846777d58607cac7b5056015a7c298d1aeb9647c0e52f1bb07b12af054a0f98c

Observation 7862238f-be87-4d00-9142-f0067f09fd5b · outbound

This paper cites Code Llama: Open Foundation Models for Code.

LLM Pretraining with Continuous Concepts Code Llama: Open Foundation Models for Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.364183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.364183Z digest=sha256:9ac08354f25d4a0bafcf8377623aaf04982f08c678c0751ea91b1bde61056729

Observation 9e5c6412-4262-4403-a520-bed940c2b047 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

LLM Pretraining with Continuous Concepts Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.369463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.369463Z digest=sha256:e3061b5e71d3a8670b67716944485b11010ab9a9d02a65c3910da52f9ca012bd

Observation 4ad03885-c99a-4063-a5b4-18148c1f35f3 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

LLM Pretraining with Continuous Concepts Gemma 2: Improving Open Language Models at a Practical Size

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.374645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.374645Z digest=sha256:4492b2ffd37b1ea6fcdda15351465f8c6482e27e1ad172d916ce44f56a8f49c4

Observation db8e37c5-f6ee-4792-a9c8-dd9307818a09 · outbound

This paper cites Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al.

LLM Pretraining with Continuous Concepts Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:54:39.969965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.379863Z digest=sha256:412ae9b9be439fc948306833c97da33e5b22a5aa2d5a8ab05af40d86777f354c

Observation 52478968-0718-47e1-8060-d0e336b4f5de · outbound

This paper cites Evaluation of ChatGPT and Microsoft Bing AI Chat Performances on Physics Exams of Vietnamese National High School Graduation Examination.

LLM Pretraining with Continuous Concepts Evaluation of ChatGPT and Microsoft Bing AI Chat Performances on Physics Exams of Vietnamese National High School Graduation Examination

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:54:39.522927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.384838Z digest=sha256:18df397396a8d80163fbb52d6f82088f74ddbf1bb2e44207908ec0e19d2914fb

Observation 549a592e-5b9d-4877-890b-469fe03712f6 · outbound

This paper cites Distilling System 2 into System 1.

LLM Pretraining with Continuous Concepts Distilling System 2 into System 1

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.389971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.389971Z digest=sha256:f58661ef94e37c476e6ef443091a42091a9358e74ca11b8e2ede6b90e7f687f9

Observation 4a1532da-f258-47d4-8e8d-40b4e0fc68e7 · outbound

This paper cites Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors.

LLM Pretraining with Continuous Concepts Transformer visualization via dictionary learning: contextualized embedding as a linear superposition of transformer factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.394799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.394799Z digest=sha256:0ba6eb3f504436ee956a99d6123b83f1d80a513eef166170bf43d5c99871838b

Observation b578b36d-6fbb-49e3-95d7-74bec819cd32 · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

LLM Pretraining with Continuous Concepts Representation Engineering: A Top-Down Approach to AI Transparency

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.400356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.400356Z digest=sha256:d75e331979ac2af1907235c0791bd78bcdbf6da960306cb255e1d19a583bf68e

Observation 3f751cd2-f452-4a7b-b04b-c249dc62e023 · outbound

This paper cites Our method and baseline both utilize the GPT-2-based Transformer architecture and tokenizer (Radford et al., 2019), with a context length of.

LLM Pretraining with Continuous Concepts Our method and baseline both utilize the GPT-2-based Transformer architecture and tokenizer (Radford et al., 2019), with a context length of

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:54:39.952345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.405653Z digest=sha256:adfe1eff545500a09b3761a4f939ed5e029896ae2bc3932177acf7b536b2647b

Observation 3fad0500-06e3-4d2d-bb8d-bbfedbd74ff9 · outbound

This paper cites Consequently, our models introduce an additional(C + Kconcept) × d activated parameters on top of the base Transformer parameters.

LLM Pretraining with Continuous Concepts Consequently, our models introduce an additional(C + Kconcept) × d activated parameters on top of the base Transformer parameters

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:54:39.915153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.416054Z digest=sha256:8784e37731d3d3c0ee401f787c14ec37cf0cef7c65f4d2159b6a46e724ac16cf

Observation 6daeb9f7-e9a3-49cd-a87a-8775bb20ce3e · outbound

This paper cites Latent concept modeling is a.

LLM Pretraining with Continuous Concepts Latent concept modeling is a

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T04:54:39.898587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.421494Z digest=sha256:8720502be2c2b12b2f7d7178ade7d188ba67e263b2a37cd444abb779f2310923

Observation fbe4797f-a9e9-4286-ab91-f0bfb5fc2ae7 · outbound

This paper cites For CoCoMix, the hidden state dimensionsd are 512, 1024, and 2028, with 8, 24, and 24 layers, respectively.

LLM Pretraining with Continuous Concepts For CoCoMix, the hidden state dimensionsd are 512, 1024, and 2028, with 8, 24, and 24 layers, respectively

Reference 1024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:54:39.933400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.411003Z digest=sha256:5a9731589dab276ca7afc9b1091861ea588206bb2a56afe1638fe10b9ace6465

Observation f7eaeea9-c1b1-47e5-a0fa-75acd4cc7612 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

LLM Pretraining with Continuous Concepts Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.342599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.342599Z digest=sha256:bfb63033b63d0ecf58eaa14eebea8a3590b7feb0b5f6377dea9eb4a247a4fba3

Observation fe0813bb-bf0c-4247-af1c-f7270a4f9eae · outbound

This paper cites OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text.

LLM Pretraining with Continuous Concepts OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.353679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.353679Z digest=sha256:7f12520f7248d272265d6d9ffc7cd4ac90ce19ee957476c64d9230d19aa32e00

Observation 551de2d2-9abc-4e3e-8fc2-a578f465e342 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

LLM Pretraining with Continuous Concepts Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.348326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.348326Z digest=sha256:83e2affceacba6c0da58063684864cead4969f4151409853755a4b1df2891663

Observation a3f613b6-363e-4383-a17e-3d6b5a4e3cc8 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

LLM Pretraining with Continuous Concepts Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.275380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.275380Z digest=sha256:b54d15d2b3d7d19db3d0ba88a192cd0771c99edfcb75ed9871159026655b2c1c

Observation 26b2ce15-01b2-4115-8fea-05981e86e29f · outbound

This paper cites A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs.

LLM Pretraining with Continuous Concepts A Little Help Goes a Long Way: Efficient LLM Training by Leveraging Small LMs

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.359172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.359172Z digest=sha256:b7a7816123a03a543b84c360253ded8cf7238cc37e52f96d86b8122189b9edd0

Observation 67942a3d-9f06-488c-b455-241cb2094f85 · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

LLM Pretraining with Continuous Concepts Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.263903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.263903Z digest=sha256:ba4f7086891d7f5a70775ec05ab88c5bcca83d0e8746bcd611f1e9e4679d14b0

Observation 6233ff7e-b64e-48c4-853d-51ba1a4faf96 · outbound

This paper cites Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.

LLM Pretraining with Continuous Concepts Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:54:39.987870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-08T04:54:39.257905Z digest=sha256:dcc5aa7889af1553e1805d40f88b63e8bf4f7eb01e489a28c550f09e71fa238c

Observation 6c35d17c-690d-4149-b86d-164ad2a0344e · outbound

This paper cites Implicit Chain of Thought Reasoning via Knowledge Distillation.

LLM Pretraining with Continuous Concepts Implicit Chain of Thought Reasoning via Knowledge Distillation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:39.286449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:54:39.286449Z digest=sha256:829038290914457a58ea056a80cf40e937812245c02a6e31e843b1c3fa62be6d

Pith citing papers

Observation 4e0d9e0a-28f1-4682-a39f-c1c5e1863568 · inbound

A Survey of Scaling in Large Language Model Reasoning cites this paper.

A Survey of Scaling in Large Language Model Reasoning LLM Pretraining with Continuous Concepts

Reference 189

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:22:08.794718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T21:20:07.238992Z digest=sha256:d8e9ba3afc40149fd84709dbe8a6f470223e2a56bb372b0d60b2e30a1d2c7484

Observation ffa160b7-62b5-4451-9cee-ba8cc6e5f20b · inbound

Efficient Pretraining Length Scaling cites this paper.

Efficient Pretraining Length Scaling LLM Pretraining with Continuous Concepts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.771507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.771507Z digest=sha256:bdb57779ed0ba3d52ce980f3ac8ac2e13cbc50d49025390518169859e6712076

Observation 1f9d6edd-9263-434b-b776-3b3437af8323 · inbound

Kinetics: Rethinking Test-Time Scaling Laws cites this paper.

Kinetics: Rethinking Test-Time Scaling Laws LLM Pretraining with Continuous Concepts

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:34.382006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:30:34.382006Z digest=sha256:2f273008085ffb13bbfdf02e9b35f15d6d5f893e55c5d790bf844d38dab20c65

Observation 3959c756-5e25-421b-b86a-7fa6cc9f9be9 · inbound

A foundation model with multi-variate parallel attention to generate neuronal activity cites this paper.

A foundation model with multi-variate parallel attention to generate neuronal activity LLM Pretraining with Continuous Concepts

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:55:46.262459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:55:46.262459Z digest=sha256:9cd1d1d00d3b46c0091e2394655944ccf16a7afcc7a8d72d8674d26d16c6304c

Observation e9f46559-0637-4498-a104-7320a4abcc83 · inbound

Towards Distributed Neural Architectures cites this paper.

Towards Distributed Neural Architectures LLM Pretraining with Continuous Concepts

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T22:14:19.808871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:14:19.808871Z digest=sha256:f7beb8b6fc9ed4d0f3900cd1efc4a67da95967e61e9e4933aaa7dcab7569e0e9

Observation 3281b352-bf5a-4ab0-83ec-ce80173c72f3 · inbound

Implicit Reasoning in Large Language Models: A Comprehensive Survey cites this paper.

Implicit Reasoning in Large Language Models: A Comprehensive Survey LLM Pretraining with Continuous Concepts

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:36.759475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:39:36.759475Z digest=sha256:b5b550e5738817155640370c51c78fb34f1dc1f5f3e696250e3f9e498104d7f1

Observation 0c4a3242-f36c-4ba6-b8b7-61579dc05215 · inbound

LaRe: Latent Refocusing for Multimodal Reasoning cites this paper.

LaRe: Latent Refocusing for Multimodal Reasoning LLM Pretraining with Continuous Concepts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T00:15:16.513242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:15:16.513242Z digest=sha256:da68089f2eb7f22fe0c39e612ed5dbad8a857534127fab9648f6ee751278fe74

Observation 49da501c-ea80-4c50-9e14-3297463d8dc9 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook LLM Pretraining with Continuous Concepts

Reference 196

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:ba4a28f63811fe17708cf7a05a9d6e15deb5c0859d04926b18bf94c204a1c5df

Observation 563aa1fc-bd0a-4c8b-ae3c-2846ce007c50 · inbound

LEPO: Latent Reasoning Policy Optimization for Large Language Models cites this paper.

LEPO: Latent Reasoning Policy Optimization for Large Language Models LLM Pretraining with Continuous Concepts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:09.869259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T05:45:50.406959Z digest=sha256:8623019af5911823009756550f99c4290fc93943730e7bf97c47d88aa60b3e2d

Observation 73909bf1-8873-46a6-ad32-f7ba0130d606 · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering LLM Pretraining with Continuous Concepts

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T23:54:45.137827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:9a848afd2ec417db80b1e09dc05251019afd0d6ded21dedc0138d2efb4058425

Observation 31659c60-6c27-4512-8158-429d649dbce5 · inbound

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG cites this paper.

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG LLM Pretraining with Continuous Concepts

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:01:12.073029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T10:27:00.257353Z digest=sha256:2ad1a95aa689e83828b92344c0b63cdd086a2edad71bee733d21b7c9663222dd

Observation 3812236f-747b-44f0-a585-107309943a5e · inbound

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning cites this paper.

CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoning LLM Pretraining with Continuous Concepts

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.793043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T05:25:51.655208Z digest=sha256:284f12ed8401a1f52eba45e561aa3361ea1ad46746afc84531d4010ff35e00e3

Observation dfdefbf8-0b3f-4d05-a234-543fdc0ceff0 · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.194388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T12:23:42.587689Z digest=sha256:52ee31dd254bd52a355fb92859c529860538213efcda194524f41ddc48f3207b

Observation 720668a7-825d-4c95-a79b-b749c11ce05c · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:16.257287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-04T00:38:33.708836Z digest=sha256:f2e81ff7a049e95d0aedaf4b394ae1721b7b20b84e9190f6aa59fffa1fa48b5c

Observation 80db826e-7060-4cb0-b06e-9e99428f908b · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:c56e3382df6410f4c87c01a8613aab349323f26cf070fce9e796202f5f2397e9

Observation 1bb8e52f-601f-4ac6-b2e2-de24ce4ad896 · inbound

Learn from your own latents and not from tokens: A sample-complexity theory cites this paper.

Learn from your own latents and not from tokens: A sample-complexity theory LLM Pretraining with Continuous Concepts

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:23:50.223099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T18:22:40.908929Z digest=sha256:a1773c4d29bb5a96a37f8bef95561c3ab6a09f30cf970bc1e6c4208bce652554

Observation b58f78ff-db56-47d7-b67e-f93805d42449 · inbound

Unlocking the Working Memory of Large Language Models for Latent Reasoning cites this paper.

Unlocking the Working Memory of Large Language Models for Latent Reasoning LLM Pretraining with Continuous Concepts

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.922213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:02:05.390732Z digest=sha256:37d2d13b758a46b9e54268a47d8a21de1f9fec78e309841806fd8762cc79ebfb

Observation 857c626d-7c49-4ea9-917c-847b4a225a69 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers LLM Pretraining with Continuous Concepts

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:32:46.627650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:f7ff6e01e0acc296059b89a0a2d8f4310c9f38345782d5219d77af69e7f2f33e

Observation ddf34b57-8542-4937-8107-2ec9681cd422 · inbound

Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories cites this paper.

Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories LLM Pretraining with Continuous Concepts

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-11T00:37:43.265957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-11T00:27:52.555374Z digest=sha256:f74cbeb036d9e64cc3257b72d05a4c373d9f342d67ba399d2ebc435d773b2955