Pith. sign in

Paper Citation Record · LEDGER

Efficient Pretraining Length Scaling

As of 18 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 4 inbound Pith citation observations for arXiv:2504.14992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.14992 v2

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.825037Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:31:30.334937Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T07:43:11.766620Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved50
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2a495a46-b68e-4c72-b08b-8ed75d565a47 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Efficient Pretraining Length Scaling Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.558018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.558018Z digest=sha256:2af896010cb4745160d0a3477ad4fe0bb911fc0ed32ef058361c685c01288a59

Observation 37f726f6-5bd4-452b-b1ee-51c5db3e5e8e · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints.

Efficient Pretraining Length Scaling Gqa: Training generalized multi-query transformer models from multi-head checkpoints

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.563405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.563405Z digest=sha256:1c161a4678112e299d8dfc800d00971c379f408436d0d521ee8bb607528f120e

Observation 15b54ffb-3f76-4735-b1f1-a3cc8cfe9928 · outbound

This paper cites Longformer: The Long-Document Transformer.

Efficient Pretraining Length Scaling Longformer: The Long-Document Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.567873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.567873Z digest=sha256:9aef199057558928a640cfdc1619538d433e07112d6be2339f1b594dae378228

Observation 36e29e45-f8e2-4088-83a9-6c5861deb028 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Efficient Pretraining Length Scaling Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.572503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.572503Z digest=sha256:1a157146d83dde8be43826f91beaf35100a1966448ab4f657348803e09053355

Observation 091c1abf-aee0-4067-8db9-777408e2aedd · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

Efficient Pretraining Length Scaling Striped Attention: Faster Ring Attention for Causal Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.576790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.576790Z digest=sha256:0c3e267a6042ecf4091f3e0048177d3c6bb6fbdad67f2782bb185c6c28397d4d

Observation f0f1e1f3-92a8-43a5-b4ed-aa5dc82d590b · outbound

This paper cites Step-level Value Preference Optimization for Mathematical Reasoning.

Efficient Pretraining Length Scaling Step-level Value Preference Optimization for Mathematical Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.581679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.581679Z digest=sha256:8b37539e82e2d8f8ccbb074289567778a5dffb9782a17c3e4b0949954eb03c31

Observation a9ecf90b-d3f6-4c55-ae68-2d61dd38abc3 · outbound

This paper cites Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking.

Efficient Pretraining Length Scaling Inner Thinking Transformer: Leveraging Dynamic Depth Scaling to Foster Adaptive Internal Thinking

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.586615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.586615Z digest=sha256:2f958fcaba1f52a918ae60952fed5ae461bb9dd0826aba595ba26c1a35e7fc46

Observation d98a2357-b202-48e5-ad30-4a0bf0bd91ea · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Efficient Pretraining Length Scaling Generating Long Sequences with Sparse Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.591194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.591194Z digest=sha256:bcdd5a6a8b0257875588bdcde94c5d90c59038f230b457c62b6440ef056f7a4d

Observation b9e293a4-de22-4678-9395-6af6dcc67457 · outbound

This paper cites Unified scaling laws for routed language models.

Efficient Pretraining Length Scaling Unified scaling laws for routed language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.595724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.595724Z digest=sha256:b5b36f9c07b283d65cabb6bf3af3a25c6025034ebcfe008fc4f9a7cabdf02d7b

Observation 80048c91-2a95-4320-867c-e50c8cc87163 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Efficient Pretraining Length Scaling Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.599886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.599886Z digest=sha256:3d2a950e3eef36df4477670a26a983002a869e29c61903dc3c695eca4fac667a

Observation 665c3ae4-19f4-4b9f-bc13-659ac7dbb42b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Efficient Pretraining Length Scaling Training Verifiers to Solve Math Word Problems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.604174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.604174Z digest=sha256:0a2fa796d56f35643937601c7d0bec6990f35047525e0c99a55b8ce5189296ae

Observation 8a8c855b-09c9-45d8-a140-9ace5b336d3f · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Efficient Pretraining Length Scaling Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.816535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.608654Z digest=sha256:8547e1a49ff8927fb7ab022466d282becb28125cdc00d2afd2cc10b5222da6c9

Observation c858dc2e-bee8-45ca-b2d5-c2fbc0795e09 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Efficient Pretraining Length Scaling Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.803359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.612738Z digest=sha256:83848f80a867b65a16259c4ec8df6c88cc685775af038ec60d81402382aa1bbc

Observation a35e5cf4-34aa-4230-be4f-26896c2c47d0 · outbound

This paper cites Flash-decoding for long-context inference, October.

Efficient Pretraining Length Scaling Flash-decoding for long-context inference, October

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.789928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.616727Z digest=sha256:014cc0413209f9697bac8a8d05286df35d68ce9f39aeb9c7172faf54acb26c66

Observation 0fcddfd9-b261-4dee-aa43-098683fe5ded · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

Efficient Pretraining Length Scaling LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.624969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.624969Z digest=sha256:752eb522f82bc3a3d2d831d16798fde1ff447bca9d04e1220f83e99e55cd438c

Observation 90fe73a9-05d6-4945-bf02-11a2bcb339ef · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Efficient Pretraining Length Scaling Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.629386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.629386Z digest=sha256:e43f2402f6e164d97b1cd63efd436bd9aba404762d651b5c8fa0308f0725e441

Observation c26e4082-807b-40ae-a037-b99d19b14464 · outbound

This paper cites Think before you speak: Training language models with pause tokens.

Efficient Pretraining Length Scaling Think before you speak: Training language models with pause tokens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.763606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.634029Z digest=sha256:31abf818c7efb089476434537d42e080d1dbebd1821f3d2aaa0a9de65fd40ca7

Observation 92849502-2711-40bf-ae22-62a82284aab2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Efficient Pretraining Length Scaling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.638548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.638548Z digest=sha256:e1d83d085cb226232758cc93dfbe03bad95fd9614746bc8f29e1ccf01475b08f

Observation 3f7478eb-756e-4cf2-aecd-d02af900b175 · outbound

This paper cites PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models.

Efficient Pretraining Length Scaling PSYDIAL: Personality-based Synthetic Dialogue Generation using Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.642741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.642741Z digest=sha256:8908e645facf9fee09649916d929b66f5d74491b4fa8c56498776875b29930e7

Observation 80155363-9798-4441-97c7-6b2949a45db3 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Efficient Pretraining Length Scaling Training Large Language Models to Reason in a Continuous Latent Space

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.646890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.646890Z digest=sha256:943412351d4cf560acf77236e55e340cf005a2a65c733501d7cbc358d0658c21

Observation 7afb1830-a59a-4eaf-b75e-33d7f4ffe78f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Efficient Pretraining Length Scaling Measuring Massive Multitask Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.651372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.651372Z digest=sha256:eac03d88fa4e6ce1c82a4125d9e81087599a637c5ebeed7ab1086cd313873f66

Observation 209e70b9-66b9-4053-bb77-0f80cd91d6c9 · outbound

This paper cites Scaling Laws for Transfer.

Efficient Pretraining Length Scaling Scaling Laws for Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.655454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.655454Z digest=sha256:fd7ae456bd30b7853b4c3468ef78732d9447b688cc5e8e3bd94395d9220b4d52

Observation f1b9f26b-e458-4b6f-9015-6ba6d282e343 · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

Efficient Pretraining Length Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.659632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.659632Z digest=sha256:de6d03b77f3509a1bfddc21b844a63c690571683e22bab338f00183ce65dd748

Observation 163c4d11-ee79-453a-af08-6709b56dad43 · outbound

This paper cites Mistral 7B.

Efficient Pretraining Length Scaling Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.664264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.664264Z digest=sha256:f570f11e1382539518ddd001c80f66242b75bb664e5129ac384c4e05e8bbe0a4

Observation ca625383-aefd-4784-9ef3-633c5428580e · outbound

This paper cites Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024.

Efficient Pretraining Length Scaling Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention.Advances in Neural Information Processing Systems, 37:52481–52515, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.669416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.669416Z digest=sha256:2516a9e69c3ceb112ca469f5203ed32f3e4e6eaa9d04e3b34f402ce60605cee9

Observation a06f4371-7320-412d-91cd-c08c3daec77f · outbound

This paper cites Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression.

Efficient Pretraining Length Scaling Longllmlingua: Accelerating and enhancing llms in long context scenarios via prompt compression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.673545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.673545Z digest=sha256:099efd759b6cb449fd08bf3b0b4e8dda272c74ceb951ae5b27e4c516ba0233fb

Observation 497d90ed-2ef1-463c-818a-3a55a1e04edd · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Efficient Pretraining Length Scaling SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.677625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.677625Z digest=sha256:91b41c6c056972f68c77aea72bce0bdbc7bf9f5c6cfd1bc4070b0c6acdee46d8

Observation 7e232f65-c2fc-48fb-afb6-a44fe20f03a4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Efficient Pretraining Length Scaling Scaling Laws for Neural Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.681973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.681973Z digest=sha256:20f93dd869d89673b73bb009dac25af1194bdcc49c410be22605983418d83aee

Observation 1a5ddb20-51a2-40f5-a640-73dd082d3a0f · outbound

This paper cites Natural questions: a benchmark for question answering research.

Efficient Pretraining Length Scaling Natural questions: a benchmark for question answering research

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.686173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.686173Z digest=sha256:631dbd0f51d12ca2dceedf1fa8c7db436fdc29260cdd972708a25312ba5ae593

Observation 3437b068-4646-4092-950c-17ae22082461 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Efficient Pretraining Length Scaling Efficient memory management for large language model serving with pagedattention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.690438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.690438Z digest=sha256:93111e4599630c11abfae55042d8c8fbad0d6e4c10808de006c1fa7f26451630

Observation a9cee53b-3d09-4738-9c19-20f8869de968 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

Efficient Pretraining Length Scaling MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.694729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.694729Z digest=sha256:f45f177cbe7331137dac1fa6d1618aae84bf12868e635a7622978b1c6e5f2670

Observation 31d8bb76-a3fe-4582-a07b-bc8dbe2f1d63 · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.698916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.698916Z digest=sha256:1b8ecf146644c071308649d9bde4640329d3572c355e9f4719ebc15a6f22e8f7

Observation 689cae58-4f74-4dae-b1f9-6d8c08e0d3ea · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Efficient Pretraining Length Scaling DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.708789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.708789Z digest=sha256:3ae7a62512ff6577f29413f38788a8134d717b6c69573c822b729b53c674d880

Observation e171f76d-316f-4c7e-ac7e-db93ce48ec98 · outbound

This paper cites DeepSeek-V3 Technical Report.

Efficient Pretraining Length Scaling DeepSeek-V3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.713440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.713440Z digest=sha256:1a240f5ce19fe0dfdfe6290d7fe3b93966c22f8356eb8826be5d86c7cd302ec3

Observation b540b48a-e3b9-4079-bcf0-63075bde9b1f · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Efficient Pretraining Length Scaling Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.717780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.717780Z digest=sha256:bd3b53a70f9aa5913f99654b2609f6af8fcd64df51a13bcb6ccc6b77472e2f6a

Observation 6970b460-588e-45a6-a150-d711cd57ba0a · outbound

This paper cites Cotformer: More tokens with attention make up for less depth.

Efficient Pretraining Length Scaling Cotformer: More tokens with attention make up for less depth

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.733667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.722872Z digest=sha256:ee85982d967949b60ce09866c82f4a7b302197214b864048d280be64cc1e8b24

Observation 10601342-b229-4e83-bc93-b1f013364f00 · outbound

This paper cites 2 OLMo 2 Furious.

Efficient Pretraining Length Scaling 2 OLMo 2 Furious

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.726822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.726822Z digest=sha256:ca11fe496d3bdd199cd8d39f70bde952c5e0ea81621eeaef74ddbae54d74aa51

Observation cf7b5a91-4723-42d9-911b-62c308fbc876 · outbound

This paper cites Learning to reason with llms, 2024.

Efficient Pretraining Length Scaling Learning to reason with llms, 2024

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.730983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.730983Z digest=sha256:2915d7df26ea44a11f782cf9738c6d7b61c106c55007cf8a323fae268e2285b5

Observation e15b3320-6f79-4d74-be8f-d37b62ddbe0a · outbound

This paper cites Learning to reason with llms, 2025.

Efficient Pretraining Length Scaling Learning to reason with llms, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.735013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.735013Z digest=sha256:37675d1952c858e12a23749e2a3c94dc1074dbc6113a89ceb052f0ca88d3f8ae

Observation 9e485649-1b75-42d9-b5bf-ab6eba06eb3f · outbound

This paper cites Vicky Zhao, Lili Qiu, and Dongmei Zhang.

Efficient Pretraining Length Scaling Vicky Zhao, Lili Qiu, and Dongmei Zhang

Reference 40

Resolution
malformed identifier
no resolver link, observed 2026-08-16T11:40:27.738918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.738918Z digest=sha256:d40d39697515c1ab1cbd50d94d857193940f46e74a306f387def234819b1e3e1

Observation a340972b-df2c-4a84-a40c-c1ea645abdda · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Efficient Pretraining Length Scaling Gpqa: A graduate-level google-proof q&a benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.742943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.742943Z digest=sha256:97c369a0fb32ace8f722c16cb55ad063e5e7bf0e65c4ac3d9c0b4118d2979520

Observation e4498814-6522-4633-b142-e43347dd63db · outbound

This paper cites SparQ Attention: Bandwidth-Efficient LLM Inference.

Efficient Pretraining Length Scaling SparQ Attention: Bandwidth-Efficient LLM Inference

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.747393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.747393Z digest=sha256:0ecf7de01e51d231bd925cdc767fbb8aca4cf47b3b330a6ce78c84d9f56efe5c

Observation 91e7e568-39a5-4459-84a7-f28fa8143952 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

Efficient Pretraining Length Scaling Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.751446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.751446Z digest=sha256:eeea31002da0c786582a0818bb80398d6c3a9d31cc6417a28b268bcdf10d5f18

Observation 942dcaf6-5bca-4dae-a22d-3cede4caf60c · outbound

This paper cites Proximal Policy Optimization Algorithms.

Efficient Pretraining Length Scaling Proximal Policy Optimization Algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.755193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.755193Z digest=sha256:72d78df6b1de629eb11c66b7733ef5bd30e6322b421c84a0fa7d0e5d8e335840

Observation 5ae0e7c7-7d5d-4826-8039-6c42aebce566 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

Efficient Pretraining Length Scaling FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.759136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.759136Z digest=sha256:d4df3a32b0d2e8c8e30c7cbac9079c6520b96322c7116464437bad15d24314d5

Observation a1ac3a2b-439b-4ba0-b529-fe90cf13ed15 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Efficient Pretraining Length Scaling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.763275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.763275Z digest=sha256:6467797b23d9515da3ab5a79b7975d0408f082afd4168a746ae17c7da71c3c97

Observation cdf1eecf-8b62-4f5e-81d5-ed42dee80707 · outbound

This paper cites Sparsebert: Rethinking the importance analysis in self-attention.

Efficient Pretraining Length Scaling Sparsebert: Rethinking the importance analysis in self-attention

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.676571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.767431Z digest=sha256:2cf1ded23c292491f0d7794c6e6ebd642694a04d250a21a7a68af23529608eca

Observation ffa160b7-62b5-4451-9cee-ba8cc6e5f20b · outbound

This paper cites LLM Pretraining with Continuous Concepts.

Efficient Pretraining Length Scaling LLM Pretraining with Continuous Concepts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.771507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.771507Z digest=sha256:df6cb22936da1a41a361a06c11937914f766667b92bdef42dbc7cd9b29d97631

Observation 9651b31e-bf8a-4581-8fe1-d93d9bb3bc07 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Efficient Pretraining Length Scaling CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.775667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.775667Z digest=sha256:aef41525502167ee30398cc353e9bcefe4cf855986bce1c5f929183cae323b37

Observation 62ff530d-9287-492f-8f52-d4e11f830858 · outbound

This paper cites Quest: Query-aware sparsity for efficient long-context llm inference.

Efficient Pretraining Length Scaling Quest: Query-aware sparsity for efficient long-context llm inference

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.663361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.780307Z digest=sha256:0bc8761f8f8ecb5d78e46371db375cec9a65571decc99afba3a89776eeec1549

Observation 9cf96c16-855e-49a5-a497-93f7bf9a3a39 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Efficient Pretraining Length Scaling Gemini: A Family of Highly Capable Multimodal Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.784477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.784477Z digest=sha256:bc546212b3b0fc9f746a0b0f104d1de38acda436813d0ca25f938d069b1ab53b

Observation f878e569-4991-4ae0-9492-22124a6c8e85 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Efficient Pretraining Length Scaling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.788641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.788641Z digest=sha256:942736bc5c09c720021e8aabfc55ef3098f406f25e567512006cd37b3de66d15

Observation d7008f2f-ee9d-4ac4-906b-d724213ddca8 · outbound

This paper cites Spatten: Efficient sparse attention architecture with cascade token and head pruning.

Efficient Pretraining Length Scaling Spatten: Efficient sparse attention architecture with cascade token and head pruning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.792728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.792728Z digest=sha256:fc004b2770c9cce3b5ebccf62ab5ae8a73b38bf528e2edbc2b79ab125cbee72c

Observation 087072a8-55d2-4a7f-9dfd-b557ca3f5767 · outbound

This paper cites Openhands: An open platform for ai software developers as generalist agents.

Efficient Pretraining Length Scaling Openhands: An open platform for ai software developers as generalist agents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.796755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.796755Z digest=sha256:f8e36f0b6714bedb444ee7ca4d2972ef44d86ab39082a2090c1be7a97be12404

Observation 506aa59f-7d7a-4a29-8cb5-61cd03e2d6d7 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022.

Efficient Pretraining Length Scaling Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.800794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.800794Z digest=sha256:611e1aac18d77dc202dc59c3ab23e97c2db0e09683e9b912506d1dbd3a1513c4

Observation b9b29f29-a0e1-4ad3-af00-201c0d2f4ad6 · outbound

This paper cites Efficient streaming language models with attention sinks.

Efficient Pretraining Length Scaling Efficient streaming language models with attention sinks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.633681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.804971Z digest=sha256:a21f1dd88f862adb666ce31abea4cd9a0191d66b51fe51698bb64e5389b3edef

Observation 742911d0-84db-4836-a71b-037ca41da3bb · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

Efficient Pretraining Length Scaling Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.808920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.808920Z digest=sha256:9f5259be317dfb9ead69c4cb1a7859231fc2c32d0624a73a77224d5b588bc2a1

Observation 2cb9823b-4fc8-4f12-bf00-ca70efdd8d24 · outbound

This paper cites Big bird: Transformers for longer sequences.

Efficient Pretraining Length Scaling Big bird: Transformers for longer sequences

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.620016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.813063Z digest=sha256:de46243e26a52f8171ee6e52bb84165ae8346ebd9cf7eb2e3f2e88e334e1169a

Observation fccd7f9e-4885-43b9-83c7-359895c4e452 · outbound

This paper cites Quiet-star: Language models can teach themselves to think before speaking.

Efficient Pretraining Length Scaling Quiet-star: Language models can teach themselves to think before speaking

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.606051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.816971Z digest=sha256:8520042a962d852c2f5870680cf15f3310ff722f5b5cdc9622583aecbfce6e76

Observation d34e39bc-7abd-4d54-8354-81e1d36388c6 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Efficient Pretraining Length Scaling HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.820980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.820980Z digest=sha256:2f8ebc3cabfd45a5695c6068831004f44eeee9828cd3596de54308d1797cf46e

Observation 41be79ca-507e-42dd-9121-b032ca2660bb · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

Efficient Pretraining Length Scaling H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.592405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.825037Z digest=sha256:f587e0d17867e1015857dbdda845e8711cf7fa06427c1139c12008f4aabc05a0

Observation 100ce4b3-c7de-4b4b-953b-e3426844ab92 · outbound

This paper cites Accessed: 2024-9-29.

Efficient Pretraining Length Scaling Accessed: 2024-9-29

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:40:28.776898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:40:27.621049Z digest=sha256:4298ec5536e76eaf4bc9363507ec68de52574a99accc6f5261bf558c5e96fa9a

Observation 2cbdfd7b-e0f9-45b7-95e4-990a0b58189b · outbound

This paper cites SnapKV: LLM Knows What You are Looking for Before Generation.

Efficient Pretraining Length Scaling SnapKV: LLM Knows What You are Looking for Before Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.704329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.704329Z digest=sha256:0e3227e4b1ae5fcc8664f3e27806e09c780be252858ffbd46e2d9506e74339e1

Pith citing papers

Observation 9e028773-dd76-486a-a7fa-06f751cc643f · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:43:11.770485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T07:43:11.620446Z digest=sha256:3d19c01e5df354c94f2bb9fdff2faf7f7cd06828be85f4911a154a7cf5aa7a4b

Observation b3c40fe4-0676-45ed-83ed-00fe7703e31f · inbound

Scaling Latent Reasoning via Looped Language Models cites this paper.

Scaling Latent Reasoning via Looped Language Models Efficient Pretraining Length Scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T07:31:30.334937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:31:30.334937Z digest=sha256:626d02245ea44870c1897a4fce53aaadd61382109cc751e2db13bfda1202d478

Observation 7003648e-fb23-457e-9306-e412ae4fdc60 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Efficient Pretraining Length Scaling

Reference 235

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:7a0d24c4b88149dadd797feddaaa32c9d6035a91c70b48b4b38e87c6d483a50f

Observation f683faf2-3189-4eec-9029-366a612255d5 · inbound

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models cites this paper.

Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models Efficient Pretraining Length Scaling

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:29.285778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:05:28.713898Z digest=sha256:9772f91267f321290a8e223c405e28d673ed2d17b3ae25582b2a365cd1c3bb54