Pith. sign in

Paper Citation Record · LEDGER

Accelerating Attention with Basis Decomposition

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2510.01718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.01718 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:54:36.489014Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73142ff2-4f01-4393-9534-36de299848ad · outbound

This paper cites Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman.

Accelerating Attention with Basis Decomposition Croci, Marcelo Gennari do Nascimento, Torsten Hoefler, and James Hensman

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.229619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.229619Z digest=sha256:e853c1fd7a8f8a4ff3bfd44e0fc0ce800eec858034ac0b0ef19db8314ce80d51

Observation 640eeb3a-151f-454f-9f2f-4da6f2fa88f4 · outbound

This paper cites Longformer: The Long-Document Transformer.

Accelerating Attention with Basis Decomposition Longformer: The Long-Document Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.269838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.269838Z digest=sha256:711484b1465f06d9f699458f169ce7b95f1b52a597d83672963c593613b79205

Observation a7045ed7-be93-4994-92ba-59ddb6e45bc7 · outbound

This paper cites Language Models are Few-Shot Learners.

Accelerating Attention with Basis Decomposition Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.322033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.322033Z digest=sha256:c6a7bf062eb208ea7d71f4d9e74d50d72bbfc7dbf84767b31d5bc100a4154779

Observation 00d8f182-ab52-47d9-8058-734c395a446f · outbound

This paper cites Linear least squares solutions by householder transformations.

Accelerating Attention with Basis Decomposition Linear least squares solutions by householder transformations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.370688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.370688Z digest=sha256:6af3242f6f60cc8ed76a68cfd745f814c984d97f1aa5ed0224a0c85874af6965

Observation 5a81a44c-5501-4bb1-a19d-95d6a2fb30ee · outbound

This paper cites u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \.

Accelerating Attention with Basis Decomposition u ker, Luisa Bentivogli, and Marcello Federico. Report on the 11th IWSLT evaluation campaign. In Marcello Federico, Sebastian St \

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.457634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.457634Z digest=sha256:c086ae900acd4b50df96c5fa6e8440a2ed363c0eddb12f13ed3c1cdde6b75307

Observation 5af29c4f-d56f-4c78-ac2c-3a8c2f8f77cb · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Accelerating Attention with Basis Decomposition Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.547368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.547368Z digest=sha256:a3849ece0a82e4c7eb18119d7ae3f42a95513fbbef8f170ec20245f29b55546c

Observation cce563b0-900f-4561-a90d-b07ee1fdde09 · outbound

This paper cites Rethinking attention with performers.

Accelerating Attention with Basis Decomposition Rethinking attention with performers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.602091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.602091Z digest=sha256:ad5677e3a18ad59fdd581c43c765b36c92b1d1c6c29abf7e23c91d9f551d320d

Observation 20be47fb-d94d-4b85-8eb5-9b30cf963f76 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Accelerating Attention with Basis Decomposition FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.663535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.663535Z digest=sha256:ee3aa968370a1296e6940c5b90e887f3e7a0cd079d36623746e8c51f2ab18b95

Observation 21a61534-ad57-4eeb-a424-c151aadd6afb · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Accelerating Attention with Basis Decomposition Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.706584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.706584Z digest=sha256:2896960aa3dcfed621385eea0246649772e2928681785771caa55df4d7378245

Observation b6941ade-a66e-46b3-b115-1816909e2f36 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Accelerating Attention with Basis Decomposition An image is worth 16x16 words: Transformers for image recognition at scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.763214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.763214Z digest=sha256:aa2994cf7dd06fad7b7488137f6fa59af1ef56f27b57b21d7564b5831211a26b

Observation a2bba1f1-1e01-4127-83a6-4f2bc1d452d3 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Accelerating Attention with Basis Decomposition Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.802792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.802792Z digest=sha256:aa6103486f9708895d98ddebe5a37da58683b85b2bd6e7172bd3d238c3666e70

Observation 6010a440-46a5-4e74-89ff-412c8e4fa407 · outbound

This paper cites OPTQ : Accurate quantization for generative pre-trained transformers.

Accelerating Attention with Basis Decomposition OPTQ : Accurate quantization for generative pre-trained transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.824068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.824068Z digest=sha256:8af81bad5d03b8820999b3ab63d4006e576ec0d08b87ddff2b53adfd1a4a49de

Observation 5a9db1f6-3352-4f02-9ef6-f5ca335451d4 · outbound

This paper cites Strategies for applying low rank decomposition to transformer-based models.

Accelerating Attention with Basis Decomposition Strategies for applying low rank decomposition to transformer-based models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.865017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.865017Z digest=sha256:a2682bb9e0cfebe2efbda9a450af39d2939da1f17b897a4a874a245838cff925

Observation a057aba3-d32e-4e15-9bf0-b17222eaee1d · outbound

This paper cites SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining.

Accelerating Attention with Basis Decomposition SLTrain : a sparse plus low-rank approach for parameter and memory efficient pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.910134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.910134Z digest=sha256:d98c44156c30cc3a57991bda968f0b2a9cecc116edd374b822015dd07e6a15dd

Observation be4e18b6-074b-42af-9377-38971a8d2709 · outbound

This paper cites Language model compression with weighted low-rank factorization.

Accelerating Attention with Basis Decomposition Language model compression with weighted low-rank factorization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:30.931321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:30.931321Z digest=sha256:cb1bb185ec7a8de2c3e5b4260aeaf94a7a1093086622d084a09e31d80d201982

Observation 2a63d98f-6da0-464d-9196-37cc203100a1 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Accelerating Attention with Basis Decomposition Lo RA : Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.029756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.029756Z digest=sha256:6a815a698c4de92f140423ff50cd21b3c4f388ebb713e1e482b5b3d7e20bb427

Observation 2b676826-9b17-4b04-b090-538dc0d1f1d4 · outbound

This paper cites From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications.

Accelerating Attention with Basis Decomposition From Low Rank Gradient Subspace Stabilization to Low-Rank Weights: Observations, Theories, and Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.091978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.091978Z digest=sha256:5bfa8b14dc45ea18fa5922a38048862914b59c5452ecb162de2d893eec0623f7

Observation a6a14013-6408-4d60-858d-e52a7e75f9e7 · outbound

This paper cites Exploring Low Rank Training of Deep Neural Networks.

Accelerating Attention with Basis Decomposition Exploring Low Rank Training of Deep Neural Networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.209617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.209617Z digest=sha256:6b499e1840c6ddd36373e7e5882c555559ba11fca2c1bc686db1d6033d866ee6

Observation 631783ef-d799-40a2-8984-3a39a3333220 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Accelerating Attention with Basis Decomposition Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.333202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.333202Z digest=sha256:b075a9f07938c2a23733d829061f89f659bbed1d675bae7e1bb558c2c373db91

Observation dedb3b3c-0143-447c-86ac-5a807bdcfb2e · outbound

This paper cites LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression.

Accelerating Attention with Basis Decomposition LORD: Low Rank Decomposition Of Monolingual Code LLMs For One-Shot Compression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.437126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.437126Z digest=sha256:a5c21462675b1253e4a8e5e90426c3fdebf76155529e3c10a19193271bfc6625

Observation ec62ac58-2159-4e00-978d-e4bd905ef6e1 · outbound

This paper cites Tenenholtz, Lester Mackey, and Nicolo Fusi.

Accelerating Attention with Basis Decomposition Tenenholtz, Lester Mackey, and Nicolo Fusi

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.557629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.557629Z digest=sha256:1b992d161735d4a21bd39957ea2648f37b5153aef1390701dd2e96d12c751c32

Observation 0f83de22-e96d-4476-9f62-7053a68070a3 · outbound

This paper cites Reformer: The efficient transformer.

Accelerating Attention with Basis Decomposition Reformer: The efficient transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.651506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.651506Z digest=sha256:9729475c7ea508fa107298ae5343d82f995526be17d02ab69ca137bbe86df289

Observation 31db1703-5b8b-4a93-b2ca-a68b152ecc44 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Accelerating Attention with Basis Decomposition Efficient memory management for large language model serving with pagedattention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.752334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.752334Z digest=sha256:b97de3920dd00336e4667ed9db68d9c194678bc6f8ba53ac05440dbf1d59351b

Observation 643ba7d0-1136-467c-b008-08c2923f039d · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Accelerating Attention with Basis Decomposition Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.894154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.894154Z digest=sha256:7519166c7e89e63d107e83abba4cc045efe47f616bc7881be246c553c7784a76

Observation 01fa899c-3f53-4895-be6e-7047fb963ed7 · outbound

This paper cites L o S parse: Structured compression of large language models based on low-rank and sparse approximation.

Accelerating Attention with Basis Decomposition L o S parse: Structured compression of large language models based on low-rank and sparse approximation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:31.978012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:31.978012Z digest=sha256:37c82d27e1cfc1cdf5ded731ebb19f2805e2f4c69510638c7e0ed247115e95c0

Observation f261d3aa-14a0-4a8d-9c5a-d529ae8ba1dc · outbound

This paper cites Relo RA : High-rank training through low-rank updates.

Accelerating Attention with Basis Decomposition Relo RA : High-rank training through low-rank updates

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.133004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.133004Z digest=sha256:ecb238375e54d27abf4be500b2fe0d434adf814f5bcfa8726bb51a4ab9169db9

Observation 77cdf235-0bfd-4f03-99a5-b557d1825d12 · outbound

This paper cites MoDeGPT: Modular Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition MoDeGPT: Modular Decomposition for Large Language Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.247901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.247901Z digest=sha256:4fd3d043afb16cd4d6b847747bd10ccac8b9a25577325ca5d004bdefcde50766

Observation 83137743-3cb1-4ff3-b11a-6cf1c74487d4 · outbound

This paper cites Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s.

Accelerating Attention with Basis Decomposition Duquant: Distributing outliers via dual transformation makes stronger quantized LLM s

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.347427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.347427Z digest=sha256:c2b7cd14108059a31bc7330e2d0ee4ab1f79a77a22f7438f3779b40b7b3afef1

Observation 9f0d11ae-4f6e-4531-a7cc-a1a6c23fc160 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.

Accelerating Attention with Basis Decomposition Awq: Activation-aware weight quantization for on-device llm compression and acceleration

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.471274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.471274Z digest=sha256:67dfabd767299ed726ac76a98b9d58465f7fd102e6dae8aabb3087e56f7b0a56

Observation 5ae5674c-ffb0-4b20-839e-613f321f3d60 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Accelerating Attention with Basis Decomposition DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.527232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.527232Z digest=sha256:51de46b9cc69b71537be408d50dd7455ae900640b8fe4f9e522fb7e5dc046cf8

Observation 07b3b64d-3b08-49e8-8f1c-d43f6e3b80c1 · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating Attention with Basis Decomposition DeepSeek-V3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.615317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.615317Z digest=sha256:851df188529c9f9f70081d413277cec2cb2bef7eb56a8dc9dbf739adffc898bd

Observation 0258c38e-9abe-4a8a-b13e-c7ebc7dac914 · outbound

This paper cites Dora: weight-decomposed low-rank adaptation.

Accelerating Attention with Basis Decomposition Dora: weight-decomposed low-rank adaptation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.698827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.698827Z digest=sha256:d047662c4379703dda4778acc1ccd4e57dec30d5514737e9444908d115fd1895

Observation dd5384e6-bb56-4bcc-8dd3-d47a66c0135b · outbound

This paper cites Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025.

Accelerating Attention with Basis Decomposition Eora: Fine-tuning-free compensation for compressed llm with eigenspace low-rank approximation, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.861188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.861188Z digest=sha256:1b358cf27be3b4d0ea9e6d5fed43663fac1f8102de8c016e6d981073f68d2874

Observation 7ac60136-ac83-4676-a145-dda4bcac9790 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

Accelerating Attention with Basis Decomposition Llm-pruner: On the structural pruning of large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:32.977760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:32.977760Z digest=sha256:708de58c957db15efd3acd93c0a7d7fd7fef28b866ed6d2bc75c4fc04b9f698f

Observation 953b5e5b-d8ec-4d75-a517-f86fd2b5d80d · outbound

This paper cites Pi SSA : Principal singular values and singular vectors adaptation of large language models.

Accelerating Attention with Basis Decomposition Pi SSA : Principal singular values and singular vectors adaptation of large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.097917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.097917Z digest=sha256:8b4c1f3cdef27b13e53c9aba8d7b94d1466d7121d18f9e7a7123f7fda356d02b

Observation 3dbea6ef-0b63-4017-9997-9c4ea74232fe · outbound

This paper cites Accelerating Sparse Deep Neural Networks.

Accelerating Attention with Basis Decomposition Accelerating Sparse Deep Neural Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.214208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.214208Z digest=sha256:46663adc4b4a1d403fd14492615dbfef718eecbcf1c01fb9a0b76e9619041081

Observation d439b2ad-6e08-4225-bfbd-55466bb6f4e4 · outbound

This paper cites Dobi-svd: Differentiable svd for llm compression and some new perspectives.

Accelerating Attention with Basis Decomposition Dobi-svd: Differentiable svd for llm compression and some new perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.322566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.322566Z digest=sha256:17fe1cd0ebacdef2606a935e2436172af6eebf534b701aa3d536a4f649c513fb

Observation 6f961544-b9b2-4dd2-8839-cc51cc6df62f · outbound

This paper cites Improving language understanding by generative pre-training.

Accelerating Attention with Basis Decomposition Improving language understanding by generative pre-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.440915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.440915Z digest=sha256:6939f2fb175fdb46259f21003138f91636ad30b5f287f5175eb069f1c649d015

Observation c43b9212-5043-4bf1-9d40-21ba7e391419 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Accelerating Attention with Basis Decomposition Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.536976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.536976Z digest=sha256:319149013563770c3f7fdf1480fd19f5523370614f846bee1c999854359a3ae2

Observation 278dc369-fde2-4308-9668-e49d0114f923 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.587104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.587104Z digest=sha256:d729be0a27d94b89d117094aa6d02255fc5584f79ae496d80feb50ffb1c5e882

Observation 31e01949-78d7-4159-beb9-0126df88ca27 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

Accelerating Attention with Basis Decomposition Compressing large language models using low rank and low precision decomposition

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.690785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.690785Z digest=sha256:70b1ee84657835c7dc44d1dda8328604feae670982a1e9cb63cfae1c102ed167

Observation 342da8c0-a4e9-4b3f-9574-f2d413e0f6b4 · outbound

This paper cites ESPACE : Dimensionality reduction of activations for model compression.

Accelerating Attention with Basis Decomposition ESPACE : Dimensionality reduction of activations for model compression

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.809317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.809317Z digest=sha256:d2366cd041fa7d55e0a1b9bc360916fd1b742097f945a8024347c52cc26f7243

Observation e920ecbc-9788-49b6-8e95-0174b00261b3 · outbound

This paper cites Robust low-rank training via approximate orthonormal constraints.

Accelerating Attention with Basis Decomposition Robust low-rank training via approximate orthonormal constraints

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:33.919849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:33.919849Z digest=sha256:527d4375706fcc7c9c02986bb4825ed1f1728fc6f3839466f712014a4c1649d3

Observation 1ea1abe1-5963-4ead-a7f4-2d53663b6ac1 · outbound

This paper cites Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations.

Accelerating Attention with Basis Decomposition Low-rank lottery tickets: finding efficient low-rank neural networks via matrix differential equations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.055161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.055161Z digest=sha256:5bab32a0cc83071d058d93e70ac8a2647730b11b31014f5128ffba551804f649

Observation 93e826c3-8a4c-4eb0-a545-a663f681d28b · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

Accelerating Attention with Basis Decomposition Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.234225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.234225Z digest=sha256:23363a5c57e218c252a969b1eb0d35e5ccf4326fdbfe800fc4ce3aba2678ee18

Observation 586d5a36-5a80-422b-b19e-72f3bccac425 · outbound

This paper cites The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction.

Accelerating Attention with Basis Decomposition The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.403786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.403786Z digest=sha256:8dad90cdd2ccfdfc880a1cfb356a0c5c707945030476f5f89058a3b2ea8856ce

Observation 528f9003-faa7-48b4-80d5-90af2d766239 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

Accelerating Attention with Basis Decomposition Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.560871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.560871Z digest=sha256:c1823066f839add10a263c9e8205ea41ff5adaf703adfc4666cda05485ea6faa

Observation 18fd2129-8fbe-4609-949c-f15b7ecfa635 · outbound

This paper cites A simple and effective pruning approach for large language models.

Accelerating Attention with Basis Decomposition A simple and effective pruning approach for large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.725960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.725960Z digest=sha256:9d96e47b4f90a1a4a641d5458dfab7e7011f6959908a977f5b57b137d3ad7ecb

Observation 3ae9a8c8-8529-4874-9453-f8d373e940e4 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating Attention with Basis Decomposition LLaMA: Open and Efficient Foundation Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.850727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.850727Z digest=sha256:f1b8f96b67add43e90d14bdf73399f474ce24b1b120690be99bfbad160c989ef

Observation 3e2ec9d2-3863-4724-a5b5-713fdfd1aa2f · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:34.978563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:34.978563Z digest=sha256:69566bb1c52955b911a6584497e3b4897169c50522695b67b157420543b0f0b2

Observation f0136548-5b07-446f-8cd5-045e713c7f8b · outbound

This paper cites Attention is all you need.

Accelerating Attention with Basis Decomposition Attention is all you need

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.063112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.063112Z digest=sha256:f93456ab19741a5f9f6147bd0856dad4a4d41656b1fc136b630a15f1631ae14c

Observation bec3bf8d-1d0f-4873-9ac4-a9fea3897be1 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Accelerating Attention with Basis Decomposition Linformer: Self-Attention with Linear Complexity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.148840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.148840Z digest=sha256:e4034433e842eb5397d7bc7c80bc7f12755e12bd69a961921834885e0edee7b2

Observation d362884c-1a04-4ac8-93cd-c40cc7ca3094 · outbound

This paper cites SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression.

Accelerating Attention with Basis Decomposition SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.258159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.258159Z digest=sha256:597b1b20f22b2e7fba0a5f0c8dd39b463ca51adf5ca1d4d6f7c7ef9c8f0b133c

Observation c321b518-a77c-447f-97dd-88711359f07e · outbound

This paper cites S mooth Q uant: Accurate and efficient post-training quantization for large language models.

Accelerating Attention with Basis Decomposition S mooth Q uant: Accurate and efficient post-training quantization for large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.361956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.361956Z digest=sha256:d9b8a736bcb50a4c483143fa40c83d8dc31f92e02a0d5a608d0660c28727c4f1

Observation 3a849161-2923-4d55-bd1f-e3a7efe8544a · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

Accelerating Attention with Basis Decomposition ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.464572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.464572Z digest=sha256:19202fd604741853e8d1063583f916d006de745eb2728e3a5d56b5fc85a52ca5

Observation fd338049-9301-4ef7-b7ea-92b6c9f09adf · outbound

This paper cites IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning.

Accelerating Attention with Basis Decomposition IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.581040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.581040Z digest=sha256:95c2bc14bdab416b03b22a3d953cc8634e51ef416100f405048b541edf1a868c

Observation a45d9535-e250-4bd4-bb31-6d366b7a860a · outbound

This paper cites Adaptive budget allocation for parameter-efficient fine-tuning.

Accelerating Attention with Basis Decomposition Adaptive budget allocation for parameter-efficient fine-tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.660902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.660902Z digest=sha256:0760ffc76717228447143e9b41b3e32ba154ad57954be24ddbdbbd4b8ff72931

Observation b1bce96b-694c-43aa-9b5e-4f059156fc26 · outbound

This paper cites OATS : Outlier-aware pruning through sparse and low rank decomposition.

Accelerating Attention with Basis Decomposition OATS : Outlier-aware pruning through sparse and low rank decomposition

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.775920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.775920Z digest=sha256:155d806663cb75550a345e09fd3671b2ef8209d320a8929ac1911e9530edbc50

Observation dc3d7516-ad5f-4dd5-a41c-de003089dfb9 · outbound

This paper cites Plug-and-play: An efficient post-training pruning method for large language models.

Accelerating Attention with Basis Decomposition Plug-and-play: An efficient post-training pruning method for large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.900097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.900097Z digest=sha256:3a72e16f30929893794ba5afa5d1df082f27d3ae079a1bee666988b16ec9dfec

Observation 0f894493-e622-4418-a1e9-097317145de2 · outbound

This paper cites Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models.

Accelerating Attention with Basis Decomposition Pivoting Factorization: A Compact Meta Low-Rank Representation of Sparsity for Efficient Inference in Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:35.977946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:35.977946Z digest=sha256:144a19fb1ddf6462ccf44c465117f0525d3af647b997737c10f8ba5950e35c1d

Observation 6dd34c8b-9224-430b-aadc-c57a7b72dc78 · outbound

This paper cites InRank: Incremental Low-Rank Learning.

Accelerating Attention with Basis Decomposition InRank: Incremental Low-Rank Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.066165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.066165Z digest=sha256:921b13037a2d28e16417b2e14cdb2284db30decc4568c65133a74a52c3e72243

Observation 2f8f7a76-881c-47a7-8042-894e98116cea · outbound

This paper cites write newline.

Accelerating Attention with Basis Decomposition write newline

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.165106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.165106Z digest=sha256:5adeb15819da9e5a07aeb073471c18a283b9309b1be2bbc0e8600ef932f7bcf7

Observation 1d17f1a5-7d81-45eb-8f50-b2065d4371e8 · outbound

This paper cites @esa (Ref.

Accelerating Attention with Basis Decomposition @esa (Ref

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.285194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.285194Z digest=sha256:e6d3abb2ef9e6354cd8812bda82a86283e1cc4fd7c83a742349ed50726a004ee

Observation 93e27143-3258-49e6-9bd8-d9078579c549 · outbound

This paper cites an unresolved cited work.

Accelerating Attention with Basis Decomposition Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.403613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.403613Z digest=sha256:fba7f146b970a90e8e302f1e8722016e0be69cfe74511dab3f6bb59e2539b003

Observation 07ee1195-c7ea-4742-86e3-af5f4901c69b · outbound

This paper cites u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2.

Accelerating Attention with Basis Decomposition u `:^!t )GeuwokcJ _ ]n?ICq .WT +BCBC &q=2

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T12:54:36.489014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:54:36.489014Z digest=sha256:07adc5572c3011fc1eb38b01d4436ddc2fa2db0233ad3b9a42b3b972d7a118da

Pith citing papers

No inbound Pith citation observations are available.