Pith. sign in

Paper Citation Record · LEDGER

Core Context Aware Transformers for Long Context Language Modeling

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2412.12465.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12465 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:11:39.748585Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7a503cc-f96d-407b-9d26-69efd1364aca · outbound

This paper cites Same as demonstrated in existing methods (Beltagy et al., 2020; Xiao et al., 2024b).

Core Context Aware Transformers for Long Context Language Modeling Same as demonstrated in existing methods (Beltagy et al., 2020; Xiao et al., 2024b)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.105767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.731601Z digest=sha256:bcb5f83d9c26e06ed5c4a905e2445617a7fe6ff844efb3735d207d4ca7081702

Observation b396aef2-ebad-4247-940c-1fdd7b936e8a · outbound

This paper cites Extending Context Window of Large Language Models via Positional Interpolation.

Core Context Aware Transformers for Long Context Language Modeling Extending Context Window of Large Language Models via Positional Interpolation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.602020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.602020Z digest=sha256:eacb91e6cc2b8b6b45b660e669ee4ea0104e675aaf6163c1b1cf93b1148da2aa

Observation f6846c3a-590f-4ab2-be44-6dd2f6cead4c · outbound

This paper cites Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers.

Core Context Aware Transformers for Long Context Language Modeling Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.611527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.611527Z digest=sha256:be873a5bb35be6c732eacd6dbbb3a6f9c9a2f9e6740083511fc2704101aba7f0

Observation 2d528daf-9b2b-4e23-8339-e90b4599b224 · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

Core Context Aware Transformers for Long Context Language Modeling LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.616699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.616699Z digest=sha256:bb180d79814c321366e1597fd6c748ce5fdcc9c5150c72cdeb37e5cfb4920755

Observation f06184fe-5287-4e72-8eca-abde9a2d60e2 · outbound

This paper cites Data Engineering for Scaling Language Models to 128K Context.

Core Context Aware Transformers for Long Context Language Modeling Data Engineering for Scaling Language Models to 128K Context

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.622030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.622030Z digest=sha256:7bcdbf01c589bc26ba98bd4b6fe49806ae078f008d3523cc91588f17defb12a4

Observation 5aa06aaf-5711-4cca-85b1-fdfbe526fe3f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Core Context Aware Transformers for Long Context Language Modeling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.626849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.626849Z digest=sha256:f825a9b825ec6342e9f15d7d26229c9cb08b799081d2eefcd54ac80fef47c374

Observation 64ff5deb-44bb-45a3-beb8-45b7e5b24b31 · outbound

This paper cites s 256 512 1024 2048 4096 PPL ↓ 2.98 2.92 2.86 2.79 2.73 Latency ↓ (ms) 457.4 460.1 461.4 462.8 473.1 Effect of Different Updating Strategies.

Core Context Aware Transformers for Long Context Language Modeling s 256 512 1024 2048 4096 PPL ↓ 2.98 2.92 2.86 2.79 2.73 Latency ↓ (ms) 457.4 460.1 461.4 462.8 473.1 Effect of Different Updating Strategies

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.070137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.743922Z digest=sha256:2e93c518a2469f53644f0219818ce34cf8b7e84892952a1c13b5513015773bb2

Observation c2fed794-ad54-46dc-85a7-6ccf3308cfd2 · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.248786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.650230Z digest=sha256:f96a104292287c6682d5abba1edddf118e4e9c1294ca01e45698d45e48932133

Observation 1601e86a-3514-4c59-8df0-9c6a8056f98f · outbound

This paper cites GPT-4 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling GPT-4 Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.654390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.654390Z digest=sha256:0e5706da376e6219964a6db2938d121ac21a9f88f77af7ba5204ea076134027e

Observation 5ff27d7e-ce5e-4b83-9e6b-e3a11d3f3680 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Core Context Aware Transformers for Long Context Language Modeling LLaMA: Open and Efficient Foundation Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.668374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.668374Z digest=sha256:6489c634a75cbb0d8c1fcc165d506ba368c18abe3077ffbeebb3688ac473a0b8

Observation 3935e5ac-b689-4170-93a0-c63a56593019 · outbound

This paper cites Qwen2.5 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.673740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.673740Z digest=sha256:cec24142be830249f068137988b830b4fb7a9a5718452af3e2d8129f292b7b04

Observation a55eb01a-948a-456f-8c96-6ffa2b6848ba · outbound

This paper cites The MMLU benchmark spans 57 diverse subjects, ranging from elementary mathematics to professional law.

Core Context Aware Transformers for Long Context Language Modeling The MMLU benchmark spans 57 diverse subjects, ranging from elementary mathematics to professional law

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.232662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.687129Z digest=sha256:b7ba3cc0d26cdf8121337ae7bea8d2902a9c4c0b4901f439dc828da866822576

Observation d37893a3-37fc-46b9-acb0-76cc17b11a98 · outbound

This paper cites base frequency.

Core Context Aware Transformers for Long Context Language Modeling base frequency

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.216143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.691363Z digest=sha256:b4fd6145eca2c1bfaae35b3f3802a46d308c18e1c466fe372a55041e1665721a

Observation 064f819b-70ea-4a43-9ac2-460d48dc0ec2 · outbound

This paper cites This enables us to integrate our CCA-Attention as a standalone, cache-friendly operator, effectively eliminating redundant computations.

Core Context Aware Transformers for Long Context Language Modeling This enables us to integrate our CCA-Attention as a standalone, cache-friendly operator, effectively eliminating redundant computations

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.198829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.697066Z digest=sha256:426a2acf7b3fcf781f016abe0c6bf91cc4b0314aedb4512662485a0400a39ed0

Observation 3518ab91-f9e6-4a8b-9662-049c403156cb · outbound

This paper cites Our experiments are based on the LLaMA-2 7B model fine-tuned on sequences of length 32K and 80K (Fu et al., 2024).

Core Context Aware Transformers for Long Context Language Modeling Our experiments are based on the LLaMA-2 7B model fine-tuned on sequences of length 32K and 80K (Fu et al., 2024)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.173335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.706475Z digest=sha256:48f567b8486a3e3306ef2dcf0d92b5e01cbca16dfc31799386bc35f256e4ea52

Observation da97c928-c70e-4f82-b2f3-acba2bf192c7 · outbound

This paper cites 16 Core Context Aware Transformers for Long Context Language Modeling C.

Core Context Aware Transformers for Long Context Language Modeling 16 Core Context Aware Transformers for Long Context Language Modeling C

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.138276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.719498Z digest=sha256:4e1409d7df93461c8c30a38d3d2a6c344f5368d0de3f6972158c216a59d4cf03

Observation f413f8ae-88d1-4e42-88b5-ddf611e7213e · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.122478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.725236Z digest=sha256:dcd058326c016fed520dedbfdb352093253be229fee0a5722165443710a31da3

Observation f7b4178f-21bc-4802-9862-1a136bc3393e · outbound

This paper cites Strategy Mean Pooling Max Pooling CCA-Attention (Ours) PPL ↓ 2.99 2.99 2.85 Effect of Group Size g.

Core Context Aware Transformers for Long Context Language Modeling Strategy Mean Pooling Max Pooling CCA-Attention (Ours) PPL ↓ 2.99 2.99 2.85 Effect of Group Size g

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.087172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.738189Z digest=sha256:8d5e7727f124289fd5b434aaf27eb188b35285705a78d62b71acf08e5f6b78d5

Observation ed662da6-82bc-4bd2-8177-f90118a375c6 · outbound

This paper cites The perplexity rapidly converges within approximately the first 100 iterations and remains stable over 1,000 iterations.

Core Context Aware Transformers for Long Context Language Modeling The perplexity rapidly converges within approximately the first 100 iterations and remains stable over 1,000 iterations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.051901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.748585Z digest=sha256:93e59558c50f4d007e34c9e72ed722aec637e8c9dc3d31142c253f2427f55c79

Observation 0e2e3f84-3068-4f03-9a30-9a7649d6668a · outbound

This paper cites an unresolved cited work.

Core Context Aware Transformers for Long Context Language Modeling Unresolved cited work

Reference 2000

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:11:40.152196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.715353Z digest=sha256:829b8bdb1ccc9da5cd2c7d590d542c67304259006e694c273815e224ccf87820

Observation 23c79c42-8f28-4294-b24e-edb9a870f05a · outbound

This paper cites DeepSeek-V3 Technical Report.

Core Context Aware Transformers for Long Context Language Modeling DeepSeek-V3 Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.644086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.644086Z digest=sha256:8eaabcd807367cce4c4c41a2fd5b350fec9b97aad5f6e23dd81b6daab8426758

Observation c43a855a-abd8-40e3-a063-3ab388aad145 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Core Context Aware Transformers for Long Context Language Modeling Retentive Network: A Successor to Transformer for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.663274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.663274Z digest=sha256:6bd1cf644ff0373db357f8d87babd5b7f5a611c4480e12289d871565ae9f32b9

Observation 970488b0-c9e3-4845-b982-5811d154f108 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Core Context Aware Transformers for Long Context Language Modeling RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.638818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.638818Z digest=sha256:f235a6cac1c269fc3942fda0c40a580054836657695d0748168a8a2b8b6a148f

Observation 6c96ee85-bc28-43e9-b57d-72b3dbe75c84 · outbound

This paper cites LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding.

Core Context Aware Transformers for Long Context Language Modeling LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.585247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.585247Z digest=sha256:5d311c562ad8d96a22d1d4be1f7db8bcf02f12a1cfcfcb50bddb4ee886b9106f

Observation 8e5a5456-13cc-4b3b-a090-19c543b30e14 · outbound

This paper cites Longformer: The Long-Document Transformer.

Core Context Aware Transformers for Long Context Language Modeling Longformer: The Long-Document Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.590791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.590791Z digest=sha256:e2232a014ca822fb6e25067c40f92b9d7cf8a2eb2e50935eb7252d535c5b4412

Observation 780698b2-f707-4bcb-a74d-c9aeed4745fe · outbound

This paper cites Chang, Y ., Wang, X., Wang, J., Wu, Y ., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y ., Ye, W., Zhang, Y ., Chang, Y ., Yu, P.

Core Context Aware Transformers for Long Context Language Modeling Chang, Y ., Wang, X., Wang, J., Wu, Y ., Yang, L., Zhu, K., Chen, H., Yi, X., Wang, C., Wang, Y ., Ye, W., Zhang, Y ., Chang, Y ., Yu, P

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:11:40.270750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T14:11:39.596635Z digest=sha256:39f2fcba431af54ecd051f08a966cb296416c5f75a74b8c0aadb17d0b102c10b

Observation 81609661-bbeb-4667-b9fa-27153bd0abb3 · outbound

This paper cites LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models.

Core Context Aware Transformers for Long Context Language Modeling LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:11:39.633204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:11:39.633204Z digest=sha256:7f42acbbf432b9418b4f96c67c41f187c10ca9502e12ae6fafa8be5ea46a38e3

Pith citing papers

No inbound Pith citation observations are available.