Pith. sign in

Paper Citation Record · LEDGER

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining

As of 18 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.12419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12419 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:13.242455Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10ecf9ea-2e27-4767-99a6-8c3b2fde81fe · outbound

This paper cites ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining ChineseWebText: Large-scale High-quality Chinese Web Text Extracted with Effective Evaluation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.103205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.103205Z digest=sha256:a2b25977732de7b4ba7b59a1d1a41d4ff277c5a6b950c3d30c5ffdcd5f33e28e

Observation 0ffdd637-eb79-4e81-9bfe-41a66ce05bf8 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Generating Long Sequences with Sparse Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.124172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.124172Z digest=sha256:4cc438e26df19ebc35a8905b2645d3dfc5a774adfe0cf025c306cb665f31a918

Observation d8e9b861-ffff-48dd-aa7c-c523e445bb0d · outbound

This paper cites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.139255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.139255Z digest=sha256:c0c5405a34c1e09e9081530252f8a5ae48efa7438f3aabd7f2039d648a124508

Observation b7c9018b-0724-4c60-876b-770efc6c2e16 · outbound

This paper cites Mixtral of Experts.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Mixtral of Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.149309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.149309Z digest=sha256:f7a8453190b1e9093fa45b1d83dd70468fb5b937ea1c80546245af802f148173

Observation f1277a8b-720f-455f-93c4-adf05d28d50b · outbound

This paper cites LM2: Large Memory Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining LM2: Large Memory Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.154209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.154209Z digest=sha256:40ec28d4b7f6dc269d18f8bb4ac97f0b7fc718d3f64a21111593d8809e394073

Observation a88d323c-1b27-44f5-9edc-13d59e00b0c0 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Cmmlu: Measuring massive multitask language understanding in chinese

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.819645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.159388Z digest=sha256:388ccb3794fb7289e23ac83040d9eb4b53bc8fd5132ef75a6e5ee891b5cbfa6e

Observation d4f6923e-5c28-47b0-88b4-46c00825abb2 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.164298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.164298Z digest=sha256:b76912d8eed902ede2522c2945995400ee1ac7a3716051022afec8bdebcf748e

Observation d674f108-fe07-4546-b148-6356289e6d4f · outbound

This paper cites YAYI 2: Multilingual Open-Source Large Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining YAYI 2: Multilingual Open-Source Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.168892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.168892Z digest=sha256:4b8a6596a10ef5f6c988afcfb22852677e0043ae647c690c0208e27116b81fbb

Observation 5bbb0316-3e14-4ae9-b64f-1085d797167f · outbound

This paper cites and Lin, S.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining and Lin, S

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.173742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.173742Z digest=sha256:a7a90c3b1f71d592890645d43fc53ccfeadb188447d14fe5a676023ed19959db

Observation fbefe8db-0d73-4a7d-bc6c-a1a9659ed2ef · outbound

This paper cites D., Man, H., Ngo, N.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining D., Man, H., Ngo, N

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.801749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.178158Z digest=sha256:1d054f3809223d3045336264c5a829b40715808522e0c334720c8560bb4aa3dc

Observation 3972f525-b459-4225-aedb-07acd82cc0a9 · outbound

This paper cites GPT-4 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining GPT-4 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.183149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.183149Z digest=sha256:7394b967838238c7a0b0dd4d0ad39b6153c2567334d5e55c2b5e2b039c2b0a9d

Observation 083e52af-51bf-4f5d-9672-fc0c05802e3e · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining MemGPT: Towards LLMs as Operating Systems

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.188096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.188096Z digest=sha256:5c7288bf59dbd85a716a74eb2006af86cd34daea41289aa3af34bec751cfd98b

Observation f5c7008d-2b31-435d-bf08-e1f8ea2e907f · outbound

This paper cites Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.192754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.192754Z digest=sha256:f31f8a4a981a17d75fca9e1862cb1c68f8aa61b10738ae70a0853b4bcf0bfcb8

Observation 1e57663f-f5c3-4323-8c67-1a5822d3e045 · outbound

This paper cites Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.197652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.197652Z digest=sha256:c3fb00f9ab9bc071318812879f3705430e1d49091b7409e116cd517b1d23d281

Observation a1af0f49-2fba-435d-94f6-2d4a7a6c8f49 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.202478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.202478Z digest=sha256:fbf29e49d5c38da22fdbcea1589b8d114a845c80eb1881c7806ad3b4287893fb

Observation bc1feda2-20a5-42ac-9a25-f641f915325e · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Skywork: A More Open Bilingual Foundation Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.207202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.207202Z digest=sha256:9ba2e4666e14fbe16ab9578f3e82075750127193c778c1b4bb278e70258e6771

Observation 7e3d6708-b9fc-4208-b9f9-f5e2fda05e47 · outbound

This paper cites Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Skywork-MoE: A Deep Dive into Training Techniques for Mixture-of-Experts Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.211915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.211915Z digest=sha256:65e44652990becad1139ae7f8e672590bfd30772ba09f0d6300fef7415695d31

Observation 9c22f5d8-e225-45e1-ba7e-79e2b7b284ce · outbound

This paper cites Qwen2.5 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.216593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.216593Z digest=sha256:e31b7794d7f7ebce45013f460834a0d2365bbc30ffbf44837513701ff34f9e23

Observation 825b28e8-4f0f-47bc-b8b7-d613da9a8ad1 · outbound

This paper cites MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining MAP-Neo: Highly Capable and Transparent Bilingual Large Language Model Series

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.221795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.221795Z digest=sha256:d6337d8b4f1f39243a313b568eeb18368f419a175eb67261dafa24df5b33643e

Observation 7d9df4fd-26f9-4f69-aff4-d974594451fb · outbound

This paper cites an unresolved cited work.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:37:13.771403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.232419Z digest=sha256:e93bf03da2afdc000f5fb72d202291f44c91a9e689351c1409e94519ad13a632

Observation a366ad2c-0efd-4a46-9fbb-8d00dbd114e0 · outbound

This paper cites It contains 7,473 training and 1,319 hand-written test questions, each requiring two to eight sequential reasoning steps.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining It contains 7,473 training and 1,319 hand-written test questions, each requiring two to eight sequential reasoning steps

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.755433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.237384Z digest=sha256:e187312dbb08687c56f354cbab042e050116f4bca4e04c4b5b6ab612338b0752

Observation caef4196-23c6-4ca0-a02d-71256b7aca81 · outbound

This paper cites Newton,” “Calculus,.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Newton,” “Calculus,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.738655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.242455Z digest=sha256:3214d5374612bde5ad37c63441cfc62e5a6814ae8b006407010628794862ccfd

Observation 6953884a-d2cc-4c52-8824-39dea6a68f0a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.134132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.134132Z digest=sha256:22284d369a9c2d2b2abd10253236a1c89570ed0d5332d75ae0963ff9280efd41

Observation 1a5b64e7-082e-4341-af5e-da75f2a443b5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.129282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.129282Z digest=sha256:7e4ff8b21e265864631b79b0b02316c5da0de43a873d2b76f3e063e9aa05d5e7

Observation 220ca2b1-b40f-44bf-a4fb-4d2a54465cd1 · outbound

This paper cites InternLM2 Technical Report.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining InternLM2 Technical Report

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.097937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.097937Z digest=sha256:f62a3a90a223f572ca47b869e8f585c3b32c780ddb343b416bb3d855beb79a8d

Observation 94cda244-ea9a-499e-a5a3-27ddf39a7204 · outbound

This paper cites ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.144294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.144294Z digest=sha256:84a6fc38705f95ddd64db04035a6e56011a9029fed5de941310feac294f14402

Observation af8ef260-a15d-4e14-91de-e2b3839592e9 · outbound

This paper cites We organize our supplementary as follows.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining We organize our supplementary as follows

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:13.786856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:37:13.226816Z digest=sha256:16cea48c63d5fca6353568bff173242a2bdf5567ef54b6e9032b32d927c95831

Observation 84fed58e-b464-4c71-8c58-4dac96563597 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Evaluating Large Language Models Trained on Code

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.108617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.108617Z digest=sha256:0719872e98c074db0446985456bd3fc3cbd544f95e110795132939e8f1af7878

Observation 8a340e0f-f967-49f9-b52a-5820bd73ee6e · outbound

This paper cites Longformer: The Long-Document Transformer.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Longformer: The Long-Document Transformer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.091654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.091654Z digest=sha256:b240d92ea28a69972ee07305342091d0226e5275bd064f626a2dbca2c973d245

Observation 40db4f9e-3cbf-42d5-a60e-9c119a3c28d1 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.113611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.113611Z digest=sha256:db2e9c5c8d35bcfeec01fa8d57f6d7e73300142039cfaa6872d758c29abf136c

Observation 6ff30059-21b8-4350-8253-f25a27c34997 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:13.118951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:13.118951Z digest=sha256:59705eef539bc7ac0a7cc7595f1e95c61b55b390081a768171a87e19e748f1aa

Pith citing papers

No inbound Pith citation observations are available.