Pith. sign in

Paper Citation Record · LEDGER

Scaling Reasoning without Attention

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2505.22425.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22425 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:13:28.835572Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:18:54.163862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T09:47:59.879253Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e310493-4137-4ef4-ae9d-4c5f6eff8019 · outbound

This paper cites GPT-4 Technical Report.

Scaling Reasoning without Attention GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.227804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.227804Z digest=sha256:e5be2ad5113143c97a1d79e5dd73b222d2ee3671b7b3ca5bb358a5252201ee2f

Observation 8cc3dbb4-c48c-4a46-a9f4-ec87798285cb · outbound

This paper cites Language models are few-shot learners.

Scaling Reasoning without Attention Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.415360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.415360Z digest=sha256:e81a945b8bdf4214e93ae92f0eddfec3db3ac111c93b75bc031f6c9f318b57f1

Observation 9fc8e080-b323-42ca-86af-59ed924e7fd5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Reasoning without Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.648195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.648195Z digest=sha256:1a7a302ad390bc91ebdfe7fe30b4153ad598e570c5d3848878dd540ecdf5e51a

Observation 69b356c1-0c92-4468-8a44-12d431ded2d3 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Scaling Reasoning without Attention OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.740105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.740105Z digest=sha256:a7fe2ac93db38e6809385bce6705228e4d5217306d9fb1a54278bf52447466fb

Observation fd5febec-8819-4651-ae3e-7af956b07aed · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Scaling Reasoning without Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.815329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.815329Z digest=sha256:8fec490a306ddae69826dce7827b607163ba0a3437c23325e24a8b487eb9870b

Observation 709c1181-f750-43e5-a30e-b624d8459e01 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Scaling Reasoning without Attention LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.947907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.947907Z digest=sha256:9b3dfb30e9b12c1f593b5f9753831c809cb300cc28a3efe4c817b6c485d7dd91

Observation c66d47db-93db-46ff-988f-a33b0611c00b · outbound

This paper cites 9 Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ra- masesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al.

Scaling Reasoning without Attention 9 Aitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer, Henryk Michalewski, Vinay Ra- masesh, Ambrose Slone, Cem Anil, Imanol Schlag, Theo Gutman-Solo, et al

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:13:30.398256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:13:27.038519Z digest=sha256:da7e713778c182c2ff6c31ac3a8c59cd6a51ea23b086176c78de43f7a1a964c8

Observation 1bab0e03-c8d6-4045-b22f-7be5514dbfe3 · outbound

This paper cites Let's Verify Step by Step.

Scaling Reasoning without Attention Let's Verify Step by Step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.137705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.137705Z digest=sha256:ccd9a8307ae241bec580449f1ec6be051c196f357a1534a27507b8b2e9906372

Observation 64cbe68b-78ec-4b6c-a9c0-6af0de2e2f80 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

Scaling Reasoning without Attention AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.227849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.227849Z digest=sha256:27b1b51ff6c19149746478fd4e81a06a7106089663f22fbcf6ad8a1a4be93ac7

Observation c9926482-f297-4aa8-b329-8cf8c955e4f4 · outbound

This paper cites s1: Simple test-time scaling.

Scaling Reasoning without Attention s1: Simple test-time scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.340471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.340471Z digest=sha256:0001ac1269e848abc28964459ba039673782052cecdd2096412d69c93a619093

Observation ab333d8c-9054-4dee-be5f-ae0b0743df95 · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Reasoning without Attention RWKV: Reinventing RNNs for the Transformer Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.469238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.469238Z digest=sha256:0311de90701b366123fa70774cd391b569ba36cd2d7eb2d23fd8a39343020893

Observation bb416504-5ae1-4418-bd1e-c5e554c7af25 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Scaling Reasoning without Attention Retentive Network: A Successor to Transformer for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.624874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.624874Z digest=sha256:18b84ef41947a1ac68face430211fe2d8a21bd50ed8bda85cce33c83e6775e94

Observation 5d3c91fc-0bb1-4550-a45f-cd0a3c0f3b01 · outbound

This paper cites Gemma 3 Technical Report.

Scaling Reasoning without Attention Gemma 3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.735999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.735999Z digest=sha256:7675af0ebd495669a0234096f5ced7eb000c58a06f6a93ef9466594527b8beba

Observation 6d6a2ed3-9257-4b80-8998-86bac1a0fbbb · outbound

This paper cites M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models.

Scaling Reasoning without Attention M1: Towards Scalable Test-Time Compute with Mamba Reasoning Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.861259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.861259Z digest=sha256:eb1bcf7a02b3b563cf538a5b67da84eb88ea6bd19a40c096127e30fac004602f

Observation 11a2ed28-ce7e-4a00-9c4d-92da4a5d57b2 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Scaling Reasoning without Attention Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:27.970162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:27.970162Z digest=sha256:4798f1170049d5749771166d7005d27b25f70c4e0b0ee354c2c21a932fbc64fc

Observation 232cce74-e538-4a71-8cb9-6ad92e78a480 · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

Scaling Reasoning without Attention Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.091568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.091568Z digest=sha256:25c05fd896d8fd348d782a47b270ff44dbf74287dd5583246f069723ed86f492

Observation 49190441-b27f-4683-bd0b-ae4a2a8fc160 · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Scaling Reasoning without Attention MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.195543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.195543Z digest=sha256:be8baf9fdd87251657eb9467e3b1ab95fd2db7797893e652aeb3c0175c97eaec

Observation cf98c9ff-b42a-496e-9dde-b54d27aff825 · outbound

This paper cites MAmmoTH2: Scaling Instructions from the Web.

Scaling Reasoning without Attention MAmmoTH2: Scaling Instructions from the Web

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.299855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.299855Z digest=sha256:8a7e334c790f4c41f3ad9aa1b656e2c30730d34bfdd26dc7dfdee7960829d718

Observation a25caafe-4ae7-4836-a253-4895d8bf899e · outbound

This paper cites SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving.

Scaling Reasoning without Attention SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:13:29.309591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T13:13:28.389604Z digest=sha256:9b80759fa61a04be636b6becfaa5a14b824c1e8fca2f946fb02dbca83f05de48

Observation fb20beab-3a9c-40c7-b91c-7a9042483b9c · outbound

This paper cites SubgoalXL: Subgoal-based Expert Learning for Theorem Proving.

Scaling Reasoning without Attention SubgoalXL: Subgoal-based Expert Learning for Theorem Proving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.536244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.536244Z digest=sha256:6b51a967fcb3c555cf2d42917cc1a16a1cb1789189dda4f64d57bd93cdd66e0d

Observation de3232c6-d290-402a-ba76-56aadcff66ee · outbound

This paper cites Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models.

Scaling Reasoning without Attention Promptcot: Synthesizing olympiad-level problems for mathematical reasoning in large language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.644116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.644116Z digest=sha256:9fe2c8228647b81c9b7ea7b02a7c984b62b3c0745d8cfc819bd102e43be05ff8

Observation 8832bc9f-8306-48f5-bb41-1d515f79f604 · outbound

This paper cites Efficient Attention via Control Variates.

Scaling Reasoning without Attention Efficient Attention via Control Variates

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.758966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.758966Z digest=sha256:54aa17c0255d1beac1e77091bf4919242c4fbc49d8f5fd13891651dd13ae98a4

Observation 069d62ff-3531-4f02-a1b6-9827a79ec468 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Scaling Reasoning without Attention Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.835572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.835572Z digest=sha256:5c136a69986e0aaba1022130d2c7aabacbf4a5835bd75e1ed0382e6a4e8c66c9

Observation 946022f1-e3eb-409b-a108-8f113b4e3add · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Scaling Reasoning without Attention Evaluating Large Language Models Trained on Code

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.496858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.496858Z digest=sha256:70817697a8bf73d0e82eb211d91bc8b2c26103c9afbcaf27d75d8b91c9aa6d9a

Observation 377417d6-7b14-47eb-95d5-5826c4681d22 · outbound

This paper cites OpenAI o1 System Card.

Scaling Reasoning without Attention OpenAI o1 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.876946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.876946Z digest=sha256:bf39df404ef94192cb39379cefdf6027c65f5a93ababbc453d1a1a8189d997ad

Observation 0379efc3-ba6b-4810-8651-41a7ca3c7fc3 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Scaling Reasoning without Attention Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:28.038685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:28.038685Z digest=sha256:a712fa4bc72bad84410d28eda61cdc3f71dc1bf54e369cf3c5709ccdd8518544

Observation f7c5e5c4-6f0a-43ba-9152-fe1d4437403c · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

Scaling Reasoning without Attention OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.292420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.292420Z digest=sha256:1d5f4cd08f5c36c0ff9923106c42d3007754540a74074285902d3138ca25348d

Observation 48ba9ba7-0e97-42a7-afe2-ebc9327e91a6 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scaling Reasoning without Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.570912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.570912Z digest=sha256:862e03b08d37ac19119ced86b286c434640ebfc30cd88f70f0937123f879d173

Observation ea777f1a-5425-4c36-a710-89fe317673d3 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

Scaling Reasoning without Attention Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:26.357993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:26.357993Z digest=sha256:2952bb3a13ad7776c8a9f9639e3f25f4f4a8150c3d39058bd6226faa1494fea9

Pith citing papers

Observation 38fd6968-c2b3-4337-9c67-ebb554f29524 · inbound

MetaLint: Easy-to-Hard Generalization for Code Linting cites this paper.

MetaLint: Easy-to-Hard Generalization for Code Linting Scaling Reasoning without Attention

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:12:02.876223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T04:07:31.283348Z digest=sha256:7e77f66934757e9c59596e0f283d0dd5bc36aa11bee702c1729fd308adc3497b

Observation 176b3fa8-1938-42be-a182-b772f314f799 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Scaling Reasoning without Attention

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:47:59.880509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:8a390300d59d1ff8a34dac357d6832b555f2d910e1996e377b562a01e4354199