Pith. sign in

Paper Citation Record · LEDGER

Scaling Context Requires Rethinking Attention

As of 14 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2507.04239.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04239 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:00:43.819726Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T08:36:31.022046Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:39:53.568674Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 806947b3-1b67-4409-ad1a-8a2be8ae87d1 · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

Scaling Context Requires Rethinking Attention Simple linear attention language models balance the recall-throughput tradeoff

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.702613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.702613Z digest=sha256:725fd9863678d67543461feca1bafe15b315cf4484b220930de52c14f55aca07

Observation 9a321be9-a37a-4f46-a65d-35fc9f223481 · outbound

This paper cites Longcrawl64: A Long-Context Natural-Language Dataset.

Scaling Context Requires Rethinking Attention Longcrawl64: A Long-Context Natural-Language Dataset

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.201678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.706036Z digest=sha256:179d4eee4e1a231c130b6114c8eed4e7ee04cb2b25d518e55e104bb88af73f31

Observation 0abaa931-dfa5-4c06-958a-3d57200f1425 · outbound

This paper cites Linear Transformers Are Faster , a.

Scaling Context Requires Rethinking Attention Linear Transformers Are Faster , a

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.192187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.708523Z digest=sha256:44b69ee86ff4f15bf3d852efbae336e18596e10c153f94f7442ea275b1fd35ba

Observation 15001e67-4366-4567-8552-779981a77fdd · outbound

This paper cites Compute-optimal Context Size , b.

Scaling Context Requires Rethinking Attention Compute-optimal Context Size , b

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.182758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.711133Z digest=sha256:1911b11e895015e5bee2a3a5f9969e468ef5db2e714ffe15e4a04f96c6a7926f

Observation d9ba248b-f5c5-4491-b746-6e7437920c18 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Scaling Context Requires Rethinking Attention Generating Long Sequences with Sparse Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.713677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.713677Z digest=sha256:629c830979a1bd91c3c20e15a1d4d09ae59f28608bb34441bc7256d2cbbc332d

Observation d5c7cf42-5e9b-4121-9364-663a5a3cd104 · outbound

This paper cites On the Properties of Neural Machine Translation: Encoder-Decoder Approaches.

Scaling Context Requires Rethinking Attention On the Properties of Neural Machine Translation: Encoder-Decoder Approaches

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.716373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.716373Z digest=sha256:4299fae1b93513984244b02313ee207fd563246f32f224ec7e25b8988642fdb1

Observation 0978e1f9-25f3-4539-9e45-37c51c68da28 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Scaling Context Requires Rethinking Attention FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.719428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.719428Z digest=sha256:652c447adb5bd5670a1510bfd84666cd1da5f2332b393a2f8aa797587e343a81

Observation 6e66f1c1-9f1e-42e9-be49-a96ac5c27d8a · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Scaling Context Requires Rethinking Attention FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.722179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.722179Z digest=sha256:c8b9c25bbe751399e8497775625266ff27684cf9dc6e2a4e34e7a7ba67eef57e

Observation ddf830bd-01d7-40a6-9941-b9507839b3f1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scaling Context Requires Rethinking Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.725247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.725247Z digest=sha256:e84d63a857afd87da7eb693877ca74b6c83404843dcdeb5b6fa156c77a12113a

Observation 69a5748b-8c6b-46e2-8da6-302d91a41462 · outbound

This paper cites Finding structure in time.

Scaling Context Requires Rethinking Attention Finding structure in time

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.173365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.728543Z digest=sha256:debc7b70aacce55a115140d4067213f018e0912ed3e34e0e50ff23d2cf67c335

Observation 7b99b5a7-1dfb-4cc0-95df-bbdf01566434 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Scaling Context Requires Rethinking Attention Gemini: A Family of Highly Capable Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.730959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.730959Z digest=sha256:0e36e26270a47965b7ea65874b24fe91fb3c5e3a2a397a6ff3b4800425f82dd0

Observation 70f9c2c1-47d3-4361-9236-70e41fe50a23 · outbound

This paper cites The Llama 3 Herd of Models.

Scaling Context Requires Rethinking Attention The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.734213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.734213Z digest=sha256:827fbac4b4cc3b8bef9ec852d53d83d42975122cd329b874e798b1cefb974f69

Observation 0b8fa29c-d77e-4367-86c1-9654208051d3 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Scaling Context Requires Rethinking Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.737347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.737347Z digest=sha256:1f333519475b46e8d01fbbe5bbfd309b5845b2bcaf4646d0dcd1dfc68bdb6c22

Observation e7dee07c-5c78-4154-b17a-4e4f653279ec · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Scaling Context Requires Rethinking Attention When Attention Sink Emerges in Language Models: An Empirical View

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.739897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.739897Z digest=sha256:bcebc8b2103f6369db53dfb9e4b23f322d0ecf7b795f38cd407e1fc14027d1cd

Observation 36c4efca-16b9-4620-97ea-a1d6f297582e · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Scaling Context Requires Rethinking Attention WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.742516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.742516Z digest=sha256:45591df106017b76e37cbad0f41c2b89368c450a64cc21ab7db013aa2b894ec6

Observation 0865f8c0-b0b2-458c-87e7-40961735e0e1 · outbound

This paper cites Long short-term memory.

Scaling Context Requires Rethinking Attention Long short-term memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.744949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.744949Z digest=sha256:7eb359835684dc80c6ade42b4466cd984d667d61e2f1008a56651f92df5ae9bb

Observation 8cd1a049-5c67-4165-8e5a-e0919b18d594 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Scaling Context Requires Rethinking Attention SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.747356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.747356Z digest=sha256:23b97fb0bdfe2847bf47948dd3899f2932ba836465910fda1acd37a79ed9acd9

Observation ef787ccc-dac0-446e-8a76-843fe5a4ab0e · outbound

This paper cites PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels.

Scaling Context Requires Rethinking Attention PolySketchFormer: Fast Transformers via Sketching Polynomial Kernels

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.750108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.750108Z digest=sha256:f484c1b1fbca519d874bb576dd649a3f4326f99d280c5de884ac1ef9e289dfff

Observation f69a656e-d9db-4e04-8f5f-c410780c7f0e · outbound

This paper cites Scaling Laws for Neural Language Models.

Scaling Context Requires Rethinking Attention Scaling Laws for Neural Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.752612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.752612Z digest=sha256:ddf0972dbd53079c31795592e06d00c0d25cfa1c36553ff57cdbaba0e32353b0

Observation 0f3f5402-67c2-40bd-8261-d2e154e12313 · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Scaling Context Requires Rethinking Attention Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.755014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.755014Z digest=sha256:ae1acbd90e2c133acf9768fb3d01a8376afbf8aaf26d4361cfcd11d244a2f7c2

Observation c9134b95-095e-43d3-aab2-45a25ed496f2 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Scaling Context Requires Rethinking Attention Jamba: A Hybrid Transformer-Mamba Language Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.757762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.757762Z digest=sha256:466dbc097dde880191499a8228bc814538a6b99313fb48ddc0329e40649ed1a2

Observation 3c25cab0-d1df-4c0d-ba57-e46784793316 · outbound

This paper cites Forgetting Transformer: Softmax Attention with a Forget Gate.

Scaling Context Requires Rethinking Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.760372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.760372Z digest=sha256:0478ef7bef43c513370b8245bafee4d453f0473a4ea2e78557e0b55f35f0e393

Observation de0c69b2-4382-4233-89f0-9de5342b9841 · outbound

This paper cites DeepSeek-V3 Technical Report.

Scaling Context Requires Rethinking Attention DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.762921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.762921Z digest=sha256:068b0aa5d57c0d57b001aca31ce433de3e0f7b3479226fd7d633a510d44429de

Observation d8782dc3-b5b3-4034-9a68-35282f763bf8 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Scaling Context Requires Rethinking Attention RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.765257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.765257Z digest=sha256:ee778f622ced68553d4096b2414f596ed113d67ee04b08529b913caf12c70849

Observation b6061d93-4618-4fd5-acb7-9e2488a55b30 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Scaling Context Requires Rethinking Attention The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.164430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.767832Z digest=sha256:420a8c4e671e34c521af714b626a15c0f66130283df21d63d8878d0db7ddaab8

Observation c3240aa8-5915-4fab-9f71-a40db2c9b1cb · outbound

This paper cites Online normalizer calculation for softmax.

Scaling Context Requires Rethinking Attention Online normalizer calculation for softmax

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.770365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.770365Z digest=sha256:3a4e34c1708d0f152abcac24263c9d1ceb94a9cfcaaf02985a3b6643e11b6cb6

Observation e3d4861c-6c2f-44bd-9932-87ae3957b2c2 · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Scaling Context Requires Rethinking Attention Pytorch: An imperative style, high-performance deep learning library

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.155385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.772867Z digest=sha256:8d18081ee802ced1b3cde6ceaf4d50dda25e5c6b14f645a4905b6b5ec934bd4b

Observation b3c052ec-0b0d-41c1-80d9-bd0172dfd71a · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

Scaling Context Requires Rethinking Attention RWKV: Reinventing RNNs for the Transformer Era

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.775305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.775305Z digest=sha256:57d74155097f8ca9d443be8dd22bdde4e21ee4e28e19c1284e12722aeb5ffe0c

Observation 9433799e-932c-4fac-b92e-dc7f46555db0 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Scaling Context Requires Rethinking Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.778200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.778200Z digest=sha256:80e481333262d1b7e2783dbd00b2347ca1a6c62c9573ad672bab1055a0772e66

Observation 4f13f900-6f37-4d09-870f-507fb57f35ae · outbound

This paper cites Language models are unsupervised multitask learners.

Scaling Context Requires Rethinking Attention Language models are unsupervised multitask learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.780672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.780672Z digest=sha256:5a784b8d240429083c4b507e59d366ef07a649705dcd4c72bb4313d75ba49712

Observation ffc2af75-1f8e-4578-b29a-6f39eeb387ac · outbound

This paper cites Theory, Analysis, and Best Practices for Sigmoid Self-Attention.

Scaling Context Requires Rethinking Attention Theory, Analysis, and Best Practices for Sigmoid Self-Attention

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.782864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.782864Z digest=sha256:bf2713c115857843d4e53fe2baa4fcc7220af71abf5d25749d6e78887e4119ad

Observation 8917eb8b-dad2-4e79-8653-4f34866760de · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Scaling Context Requires Rethinking Attention Toolformer: Language models can teach themselves to use tools

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.785494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.785494Z digest=sha256:074d735b3a2b819c4be9e05e6ccf50f87d6187c579016f39a9d766cefe729c80

Observation cc5957b9-fae7-4ad7-9188-88264c360130 · outbound

This paper cites Linear Transformers Are Secretly Fast Weight Programmers.

Scaling Context Requires Rethinking Attention Linear Transformers Are Secretly Fast Weight Programmers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.787796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.787796Z digest=sha256:87864036a4954c977fcc03535437c5c53a877a4521617e1aebb0861fe1b67164

Observation 340a5dad-f0ea-4d3d-8b60-55d04d75727c · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Scaling Context Requires Rethinking Attention Fast Transformer Decoding: One Write-Head is All You Need

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.790462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.790462Z digest=sha256:7b4ebd6d2567e187230b4ce562e3dd40861d9a718d63cb748b5f0411b7e1e382

Observation 99fa56f1-5187-4570-87fb-3d771cdb7fcd · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Scaling Context Requires Rethinking Attention Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.792979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.792979Z digest=sha256:191539a1484d15542e24d1b5f7f3adc852be6abd7cebc9e0748975bf141441ba

Observation 31b0ec1e-54f6-4a0e-b189-af035fd6a35f · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Scaling Context Requires Rethinking Attention Roformer: Enhanced transformer with rotary position embedding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.795421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.795421Z digest=sha256:bfa69f812214272be16eb3fde0d4444744c274f50ccba1434b3d26426884718e

Observation 271a02ea-2e44-48e7-b63a-15cb5c0ef046 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Scaling Context Requires Rethinking Attention Retentive Network: A Successor to Transformer for Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.797989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.797989Z digest=sha256:11859af82770c2ecb7213f2e6fc65be9c0d90b5b50372df07f61554318badb73

Observation a090348f-fb49-4bf1-b2fe-8c8226c7b231 · outbound

This paper cites Triton: an intermediate language and compiler for tiled neural network computations.

Scaling Context Requires Rethinking Attention Triton: an intermediate language and compiler for tiled neural network computations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.130419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.800486Z digest=sha256:b9c951c861bdd4f94b7b19ccc13fae060ee6424cb0f8ac8e90b68bbd183c884a

Observation 5ce6fe30-dd38-4ab1-9ac0-c52e11ed2101 · outbound

This paper cites Attention Is All You Need.

Scaling Context Requires Rethinking Attention Attention Is All You Need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.802813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.802813Z digest=sha256:8d66cf055a9daf85da117bd4bd7758f59c2ddc7b524c66d7cbc88fb4242fcb0a

Observation c7f95000-2018-4ad4-976e-9ed195be53c4 · outbound

This paper cites Deep neural network based low-latency speech separation with asymmetric analysis-synthesis window pair.

Scaling Context Requires Rethinking Attention Deep neural network based low-latency speech separation with asymmetric analysis-synthesis window pair

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.121345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.805156Z digest=sha256:6c5a6c7121482f8dc2e9058716473b702a8507eea11c3ce598cb4bc1b2f17f15

Observation 01cf6309-4cfe-4ee4-aa35-ff0de4a903c9 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Scaling Context Requires Rethinking Attention Chain-of-thought prompting elicits reasoning in large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.807571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.807571Z digest=sha256:b032827bcbf17dd7efc925c82826346f6383b580854530f97c4ee9bf12193ffb

Observation d4c99c64-0170-4b19-822c-e4d807b1f218 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.

Scaling Context Requires Rethinking Attention Swe-agent: Agent-computer interfaces enable automated software engineering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.809995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.809995Z digest=sha256:b5e8f18a35b6b4a097367ae2dfbee805ebb93f229a52d9fe4fc3dd417646870f

Observation e780f477-85c6-49b7-b086-b402b9413383 · outbound

This paper cites FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism , January 2024.

Scaling Context Requires Rethinking Attention FLA: A Triton-Based Library for Hardware-Efficient Implementations of Linear Attention Mechanism , January 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.101560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.812248Z digest=sha256:b4f243a072c869f273d17e3ac421106bf616c404ca95b0e8d3103c11c66ebb35

Observation 52bdd85c-618a-4463-b95f-51a7ae527cbe · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

Scaling Context Requires Rethinking Attention Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.814522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.814522Z digest=sha256:be3f782e402d1bb05cb12b5984509061063ecafff471edadcda1bde822538c8a

Observation 1fb8bbf4-257c-4440-a4e7-8e26f39e5bfc · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

Scaling Context Requires Rethinking Attention Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.817075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.817075Z digest=sha256:2eab0cb1c4b32a5ce8357594519cf2652a4ee5be91cd28141785a63dc0a6d0f5

Observation 6231b891-067a-4b46-bc12-2ac6d91a719a · outbound

This paper cites Gated slot attention for efficient linear-time sequence modeling.

Scaling Context Requires Rethinking Attention Gated slot attention for efficient linear-time sequence modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:00:44.091757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T20:00:43.819726Z digest=sha256:a691ba927fffc3a875515ef7b7963d52887addcce05188c97c3e9e0e1cb1e69b

Pith citing papers

Observation 2b7fe290-ff18-4b06-b3a3-4ec28343dc6c · inbound

MeMo: Memory as a Model cites this paper.

MeMo: Memory as a Model Scaling Context Requires Rethinking Attention

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T03:19:43.560611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T03:17:23.202604Z digest=sha256:d1d9b7cabab0cc119a1c7d7120c2a109a2e01e45d2aab7f175e5ce992e07e8b6

Observation bad90f9a-c9f3-41fc-b6f0-ac00426ca8ef · inbound

MeMo: Memory as a Model cites this paper.

MeMo: Memory as a Model Scaling Context Requires Rethinking Attention

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.570748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T08:36:31.022046Z digest=sha256:b32ab0746b70874a52b7762a0f7a959ad9ce16bffe0cfa77b39e33d540288494