Pith. sign in

Paper Citation Record · LEDGER

Multi-matrix Factorization Attention

As of 20 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2412.19255.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.19255 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:55:04.061512Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:52.656467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:32:36.395278Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bcb67969-c01d-48d7-81b1-b21ff41c4f52 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Multi-matrix Factorization Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.742782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.742782Z digest=sha256:77970c646ba8d5449c7c958b9040c95045997e956910343a76c22770419186ce

Observation e76f2303-2042-4701-89b5-4dd16375b2e6 · outbound

This paper cites The Falcon Series of Open Language Models.

Multi-matrix Factorization Attention The Falcon Series of Open Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.749966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.749966Z digest=sha256:130152edb9ebfe15310ab92fca981f3ed2f81fe35748d3e10e89fc6dac371aa3

Observation 3caed963-bb3b-451a-b038-868ea514c13b · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T00:55:05.201098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.756583Z digest=sha256:ce9a3ae2a5df11523259600d457a21b31d95ef73c6a4317c145636af7bf0256a

Observation 92b67c6b-c2ad-4f6e-80ab-be3fd0fa0894 · outbound

This paper cites Low-Rank Bottleneck in Multi-head Attention Models.

Multi-matrix Factorization Attention Low-Rank Bottleneck in Multi-head Attention Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:55:04.965410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.763727Z digest=sha256:e1d5ce411888ec9a33dec46647946a960b893caf045122ec2825102edf6f33d9

Observation ae81c9a7-50c7-4072-a22b-4d00f3febee7 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.772492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.772492Z digest=sha256:08c10ebc823200cb4229c6cecbbf9989e38bb79b61a8aabba08aa422028e9661

Observation e6d12d0e-1a79-4b0e-8f87-f82d9fe68267 · outbound

This paper cites Reducing Transformer Key-Value Cache Size with Cross-Layer Attention.

Multi-matrix Factorization Attention Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.777756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.777756Z digest=sha256:a5d6e38b6a090c8d8080761e687b5d69d420ab899f40ad54e3a1f25aeb987cbf

Observation 678b5c3c-087f-4b4a-9bef-24fa8d8c3f0b · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Multi-matrix Factorization Attention BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.784450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.784450Z digest=sha256:b2a48a819c4499098752abaee43c62542771fb2fe5e84d9cca91c4303efbc5e6

Observation c451885a-bfe1-4cfe-b4a3-0e86e51439ff · outbound

This paper cites Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y.

Multi-matrix Factorization Attention Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.789651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.789651Z digest=sha256:609733f60b00220a86237a8d7b706eb2a9853ed97a9e15348a68c5a1b25cb283

Observation 73001b93-6c2f-4a70-9f99-d319f3a44414 · outbound

This paper cites FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.

Multi-matrix Factorization Attention FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.794878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.794878Z digest=sha256:91c264a0a8dbbe8a6d1d1e5dc8a2a5fcf661c3ccbf1993524c9ce59075f4c7f4

Observation db60e678-aaca-40c2-8c99-d91446af1d55 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Multi-matrix Factorization Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.802359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.802359Z digest=sha256:cde32a420ea94a13319408c03b3f55aab581a710fe527a2b161d39083165a39b

Observation 7d058330-4dd1-409d-9251-92e2e7dfc3b0 · outbound

This paper cites Fewer Truncations Improve Language Modeling.

Multi-matrix Factorization Attention Fewer Truncations Improve Language Modeling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.808027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.808027Z digest=sha256:3b0773b418015e920bb50748a034b6e1119a8fa00dca1c59cf6a706d75041e47

Observation 8f0ae81e-355c-4074-a3a8-7d89bebb67d5 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-matrix Factorization Attention The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.812973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.812973Z digest=sha256:ed1cae488cf4469839026a606a001d130af943e33f00dcea8fb08f57038978e7

Observation 27d5b3ed-c61a-459a-8222-f31e5773b378 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.819178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.819178Z digest=sha256:5bbce8c7526a9c2320786d0a7c95d0f0ebc316644cbb905c3ac1e698cfc95080

Observation 16a0bccf-9d73-4f53-abfc-ae3a4467ff13 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Multi-matrix Factorization Attention Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.824319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.824319Z digest=sha256:d1cd4ce42a833b4728f6fce8bdc01c9de2b2b0990228d2606ba1f9a06c93cc51

Observation 9e1f3c55-eea6-43c0-ab90-debb06b7fe67 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Multi-matrix Factorization Attention Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.828877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.828877Z digest=sha256:a814534b099a7482126f410f0e5ea4bd1e5d7c3111434a3cd95865ab780ada96

Observation 515c7370-7777-40ac-8362-1158ad466c85 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Multi-matrix Factorization Attention Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.833918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.833918Z digest=sha256:5f31b10996d7f721e1e3e14c4576be509c167fd75e77773593d77c9d1f98127a

Observation 1e29b1d1-8265-487e-b149-d23950b0e44b · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

Multi-matrix Factorization Attention RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.839350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.839350Z digest=sha256:031b4386a416f7e0efb500c2eab9e329b4bcfcf227be5dcc8464f9b4c9521629

Observation 8fcc3c4c-8363-4af5-8151-9b398e34449c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-matrix Factorization Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.844509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.844509Z digest=sha256:da43b171e90fd4c70b5995b77a64cbd7662af15095e6d146598900ba8b3ecfa3

Observation a5f8428a-9483-4257-9547-65091a61dd79 · outbound

This paper cites Mixtral of Experts.

Multi-matrix Factorization Attention Mixtral of Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.850264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.850264Z digest=sha256:03246c703c8dcb24f81243853c1fecd3026c26b698d301151aa693a0e6d60182

Observation e29ba3a1-f6a0-474b-a947-cb21c7a285ae · outbound

This paper cites Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention.

Multi-matrix Factorization Attention Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.855619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.855619Z digest=sha256:4e1024837648f1278ad8a3a20e2d6cfc9bef566c700bdfdc0a00eec12fc319ac

Observation 26de0459-1aed-4a8c-9acb-590f0067ebeb · outbound

This paper cites Weight decay induces low-rank attention layers.

Multi-matrix Factorization Attention Weight decay induces low-rank attention layers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.863080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.863080Z digest=sha256:6be5e77ef9f091ab0a8785a156432a9a3e2c1db95d5f19c5f63fd1fa582ec78f

Observation 0e627618-f2cf-411c-9900-7e6e3932e4b3 · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

Multi-matrix Factorization Attention DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.868518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.868518Z digest=sha256:81900a8eb9c050013f302dd013dbfd3838a44a9db85a18a6fdaa6cf9a5e3e091

Observation 2d07eded-e829-4e6b-8ea6-878072c60e98 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Multi-matrix Factorization Attention Jamba: A Hybrid Transformer-Mamba Language Model

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.873620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.873620Z digest=sha256:a855548ec38490a554f227e514410286f95074e9fb6def206529b85023a3dc67

Observation 7f855f8d-2a7d-47cc-962b-b53268c255a9 · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-matrix Factorization Attention Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.878675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.878675Z digest=sha256:f986607f5181be5c18f04ad3f0c26382c1b71043e4786991e31e5c120804431b

Observation 2a33144a-e033-44d2-8b79-e10b0888ef76 · outbound

This paper cites Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention.

Multi-matrix Factorization Attention Scalable Efficient Training of Large Language Models with Low-dimensional Projected Attention

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-11T00:55:04.155091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T00:55:03.884161Z digest=sha256:88525a5eb6a5188574df6e9906e98d0c5f511e017f4eb561cf4a90cdb1dff601

Observation 3e62234d-b87f-4364-82fe-676556fbf698 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.890189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.890189Z digest=sha256:6d10868f268ad7844a9a4684f42bed022e3bb2d28935a81ab66943abe13525b2

Observation bd718645-5319-4ee1-8220-a2a76148f831 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

Multi-matrix Factorization Attention OLMoE: Open Mixture-of-Experts Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.895257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.895257Z digest=sha256:d290a7aa9b1d364d8d3725ff8550eaa103e792c26b733d769675b7620b6cab74

Observation 52a26e0c-be81-4f2f-9c4c-ef851eb31be0 · outbound

This paper cites Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM.

Multi-matrix Factorization Attention Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.901462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.901462Z digest=sha256:34b094e2915824741a8cbaa78eb672663734641a557b65829e256de8046bd439

Observation f9372660-e69a-4365-9c66-3cd35823a340 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

Multi-matrix Factorization Attention Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.907073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.907073Z digest=sha256:cf21392c17eb90efd6479bc7cc0e1109925402650b4526103c973d9efefa08af

Observation eab54852-8589-4d4c-996e-0ade451bf39a · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.912598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.912598Z digest=sha256:38e36598fb6a96dc66e2962ecc00a7d8a9143c43f62d5345b5f1af4a8ae50be7

Observation 2439adca-13bd-43e4-bbc1-ec73fe15d66b · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Multi-matrix Factorization Attention Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.918833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.918833Z digest=sha256:539527bb40471c9e83cb4efcdaca5587ac964669550ae843efc3b4784b257d14

Observation 5a745ef7-1ac7-4647-81af-de117a6f7d03 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.924736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.924736Z digest=sha256:a4519f3cd5cb2b959b80e3450e1d750c03f946691f98702b783fbfee4534fa6d

Observation e845b31c-b6c2-492a-acbb-35a364e79dd2 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.930221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.930221Z digest=sha256:dc4d9620f36e372eb35d675e0ff652933b88cc62ca82f060015b627d6b13ac8b

Observation c53b8bad-ebfd-4a94-b39d-9071aada4429 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Multi-matrix Factorization Attention SocialIQA: Commonsense Reasoning about Social Interactions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.944483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.944483Z digest=sha256:3d21cb9584ee700338f717f293fa614bac2dbe37f5a20536b7dc6223989b3680

Observation 044f62ea-03b4-499f-aa1f-934571b1e6e5 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Multi-matrix Factorization Attention Neural Machine Translation of Rare Words with Subword Units

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.953507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.953507Z digest=sha256:52f837c9bb12f264fe1315a23e33e7b751c719ec84e6cac8123955b7d9cdf934

Observation 9dc26d3b-df67-433d-a6db-e221fa417004 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Multi-matrix Factorization Attention Fast Transformer Decoding: One Write-Head is All You Need

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.962019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.962019Z digest=sha256:a952a89fd8489076268fba298fa32e85e2e6d7877ffba64476803a03433e8c67

Observation 46481ceb-e592-43b7-953d-e3d8f80a6bb7 · outbound

This paper cites GLU Variants Improve Transformer.

Multi-matrix Factorization Attention GLU Variants Improve Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.969065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.969065Z digest=sha256:ab1a907e6e9fd197aba81b36923f1adbc7d08bdc0aabb4b5d2e2bd2889a4f924

Observation f0cebf4b-e2e1-47b6-830e-0458faf94f32 · outbound

This paper cites Talking-Heads Attention.

Multi-matrix Factorization Attention Talking-Heads Attention

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.976160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.976160Z digest=sha256:04cdb70df4a7bb60b362f64cf8a23b430aa935fd11e43e57b2c9161aca6ee6b9

Observation d599db10-43d2-418a-9ce6-335e5f5feeeb · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Multi-matrix Factorization Attention Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.982953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.982953Z digest=sha256:f13c3f6babab373f4970442f869d1902a607f26e16c8d9c6a8bdc624d355d9d0

Observation d1ad1292-b006-46c8-914a-c6e9b46962a9 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:03.989969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:03.989969Z digest=sha256:d651c7d1efddec7012eee1ba344a71ae8fdb239f95b12dda0893932dabb0f507

Observation 36bc1ab2-a1db-49a9-a90f-bcb34ab95f26 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Multi-matrix Factorization Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.003033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.003033Z digest=sha256:5573ed1e5540e082df5260bce5a0ccbb4052ca4d2f7459f94d12cb618da2b680

Observation ecd3b455-2d7e-4233-af32-645607228eeb · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Multi-matrix Factorization Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.009622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.009622Z digest=sha256:6f1ccff50d6526a0108bf62fe1a02ccebada46a66be47c4d5717c23349c41bce

Observation 3ca8e6cc-aa11-482a-a0b4-51a81a288680 · outbound

This paper cites Attention Is All You Need.

Multi-matrix Factorization Attention Attention Is All You Need

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.016132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.016132Z digest=sha256:85ad71c1c8bf0be4a5bb1996d1fbe4d0537fd9b90e9ddd19a520d8de96eb178c

Observation 3dbcce2b-91ea-470b-b348-b6e9d2e20cf3 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Multi-matrix Factorization Attention Crowdsourcing Multiple Choice Science Questions

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.021750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.021750Z digest=sha256:45e49fcbf9a2f25f0a99b0185c88adb5d01ebe8dd11b411b5760081c6c6dff18

Observation c02ce927-b42e-47cb-8078-db7de2b8cb3d · outbound

This paper cites Improving Transformers with Dynamically Composable Multi-Head Attention.

Multi-matrix Factorization Attention Improving Transformers with Dynamically Composable Multi-Head Attention

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.027465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.027465Z digest=sha256:fe39b3ffb8b4e4c4f8995ca0216f453393c6057bccb718257a3be211b4ab9c9f

Observation aa6f7a4b-58ba-466b-98c5-7c236e74f2fd · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

Multi-matrix Factorization Attention LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.032759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.032759Z digest=sha256:542ad134d7d5108fc4fffc42c48eb0107de039f6856dca9a9f65bf1ab1128365

Observation 89770ec0-a618-4696-b185-7fcf6e6be6f2 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Multi-matrix Factorization Attention HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.038913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.038913Z digest=sha256:1af6301e64af38963ac7d6f0cd0a7038b5e027bf6b8a2a69dbbd06bd7147054d

Observation eccf9207-bffa-43db-88ba-eb1282551107 · outbound

This paper cites an unresolved cited work.

Multi-matrix Factorization Attention Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.044373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.044373Z digest=sha256:97868ff7c6264727aa3f2542b5046827542c6cd33be61cfb7a05fd2f8554a589

Observation f6a75ec7-d879-4c79-ab6d-85c446259519 · outbound

This paper cites MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding.

Multi-matrix Factorization Attention MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.049741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.049741Z digest=sha256:08f1e87ecb20036add7a2f9cbc3ef74cf7392f91402aeada5748c7962e2bed38

Observation 311f6198-0f32-4797-bba9-abeef5f0edbf · outbound

This paper cites online" 'onlinestring :=.

Multi-matrix Factorization Attention online" 'onlinestring :=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.055304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.055304Z digest=sha256:bbb730661cd50ac49d36f7761c8e0b34daab545c774b1635a4e44d50eb1d8662

Observation 647f227c-efdd-401b-adef-4d3c79a9398c · outbound

This paper cites write newline.

Multi-matrix Factorization Attention write newline

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.061512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.061512Z digest=sha256:892ae17fa065945dc49f150ba6fbf3d656b29812528c16a139a41abc37ed3a7a

Pith citing papers

Observation 8d6772a9-cfea-4d8e-8201-6ebe13e2042e · inbound

UHD Image Dehazing via anDehazeFormer with Atmospheric-aware KV Cache cites this paper.

UHD Image Dehazing via anDehazeFormer with Atmospheric-aware KV Cache Multi-matrix Factorization Attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.656467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.656467Z digest=sha256:14b39452d6088511f775e6d8fe5d541b190fba48cf4dc80d03a6ef033f8cb8d2

Observation 52e12996-15d3-485f-bf64-f646d8edfe21 · inbound

Hardware-Efficient Attention for Fast Decoding cites this paper.

Hardware-Efficient Attention for Fast Decoding Multi-matrix Factorization Attention

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:32:36.488542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T13:32:30.968649Z digest=sha256:218c3c848ebac795969d42084b5dfb89ae0cdb27432522a458dd9f2cd9170d3a