Pith. sign in

Paper Citation Record · LEDGER

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

As of 18 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 100 inbound Pith citation observations for arXiv:2402.19427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.19427 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T06:58:17.370396Z

measured 140 of 140 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 100 of 109 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:30.286857Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact34
  • verified fuzzy4
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 87cf6716-212c-4ad6-9893-b3cc7fb964fc · outbound

This paper cites GPT-4 Technical Report.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.404105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:4ef75383c21172bbce54af2310ff1e8f251104f8c93fd1de2a417687dd4ecaae

Observation 09c0fc6e-96a3-46af-bd4f-7379ef746ff6 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Neural Machine Translation by Jointly Learning to Align and Translate

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.411532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:9c89f0a79b908f66bb3b4d4c8a7684afc0a89e3740d42b370d41ae6eef94a94e

Observation 3159a262-5e7d-4778-98d0-e2f5f701be8f · outbound

This paper cites Longformer: The Long-Document Transformer.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Longformer: The Long-Document Transformer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.418193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:b5fb0a42668a97e8de2ff68c173a1fbe4daab2836542b52ba6f72862c7c90b1b

Observation 1eec6377-7683-443d-900a-c9ea21c084a4 · outbound

This paper cites Quasi-Recurrent Neural Networks.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Quasi-Recurrent Neural Networks

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.423878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:d5f3f8dd2944631b4450646ed58ecda5019a01f07e9451b7b820e09d4fdc4caf

Observation 2318a795-c3f2-4416-a960-ba6f7c47c756 · outbound

This paper cites T.Brown,B.Mann,N.Ryder,M.Subbiah,J.D.Kaplan,P.Dhariwal,A.Neelakantan,P.Shyam,G.Sastry, A.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models T.Brown,B.Mann,N.Ryder,M.Subbiah,J.D.Kaplan,P.Dhariwal,A.Neelakantan,P.Shyam,G.Sastry, A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:58:17.614165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f24160c72f101678183c9ede00c429c92c89decc79cbcb04ed518ac44ad99409

Observation c33bcbbb-ef9b-4a3e-9f87-9d6a57a05915 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.469338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f741140b797b410eb8e4ddefd6cea5d2fe5587329acfb927e24b930cb7466d22

Observation 55ee1edb-eb88-4c0f-8c5f-98566dd9df0f · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.474817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:69d719fa3f69ac47edc1034a527561e44358e84cb297fec0f98d939e0fe16ac0

Observation 5b5c40d6-e234-4fdd-b727-9b7e275ac8ae · outbound

This paper cites Hungry Hungry Hippos: Towards Language Modeling with State Space Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Hungry Hungry Hippos: Towards Language Modeling with State Space Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.482210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:59f6bbb662b8dd64c18ff59168e08f92be09705464794cfe81a7f03f592f1e1e

Observation 6e4486d7-ce5c-4c22-a8db-ec209bdd790c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.488268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:ecca911a62aca745ab1761f020f0e5c812cc918341499df732706d13cb0deed5

Observation e19687ef-b041-4cc8-a9c0-99459888b257 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.494676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:46ad903c515a2c19d0cfaa3b40dd273712bd0578499fc0e8c0136b76fb252db3

Observation 4ff28c73-b19e-4a42-946a-cbcbaebd3c2c · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Efficiently Modeling Long Sequences with Structured State Spaces

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.501302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:4a27b20b2b200f11483363461d1eb9e830769fc7deec45e7bcb7b4f04b62f3ca

Observation 0d3ed0ee-309d-4b95-afe4-dfd78e21c628 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Gaussian Error Linear Units (GELUs)

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.507685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:0e2379a672bf22234b48f997cef544aeb60bd09c391803cf1a7d078f9b6e988f

Observation 88e56c38-da17-4c57-9681-d19afd23902e · outbound

This paper cites Training Compute-Optimal Large Language Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Training Compute-Optimal Large Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.513427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:b6b4d8f73edee9b02ed2779beaabd4b54cc6b8ea3bb4462dd1b9c884309920c1

Observation 4785e0e2-ffce-4e15-87fe-1fa1c43050d1 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.519265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:48a60a1b110a5f4745bdc6e5a62118b1308e14711c503ecacc875a54856bafac

Observation df4db80d-37ac-4ded-a8b4-632fe1c177d5 · outbound

This paper cites Mistral 7B.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Mistral 7B

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.525125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f892b68e21e7a475cef99401066b74cb1bd0ca1abaebb165efbd78646b238b29

Observation 4d88dc55-0d64-41e8-97b3-4810ca6b7ef9 · outbound

This paper cites an unresolved cited work.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:58:17.617708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:a317fb3012663883f08a69248376db96bf625181e8689f9e133bf49bbd53eba7

Observation 79320dd5-00a0-45de-9a30-1e04cd95b9ee · outbound

This paper cites Scaling Laws for Neural Language Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Scaling Laws for Neural Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.532642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:c56d4619a0a33650839fa19aeaa0894475c5f60cc7218372c20e10d1d76855f5

Observation 2e74e5b8-18d7-4c59-85a4-db39883f2d70 · outbound

This paper cites GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models GateLoop: Fully Data-Controlled Linear Recurrence for Sequence Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.539189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:0e663d94a3831aca17635773538872cc3ca1189ec3d19553e8a72208ef12dbb1

Observation d642502c-c943-4d42-a12d-22b4c622f805 · outbound

This paper cites Advances in Neural Information Processing Systems,36.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Advances in Neural Information Processing Systems,36

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:58:17.629689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:b8d44c58ce731cc1a133e4226d09b566849a19222e82bbbe259d7333f3233937

Observation 92df91aa-32dc-4d4b-ab6a-87b2f27590c3 · outbound

This paper cites Decoupled Weight Decay Regularization.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Decoupled Weight Decay Regularization

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.544914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:5419274c141546a576c5b162458e7667d44e94398082b1aeb992c012bab8a189

Observation 5faca210-37cd-46fe-b80f-67a2fe77813a · outbound

This paper cites Parallelizing Linear Recurrent Neural Nets Over Sequence Length.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Parallelizing Linear Recurrent Neural Nets Over Sequence Length

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.550241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f1759d5fcd370cba974413f759ba6e656f586b2eaf808f0b5e1eafce834f1ce6

Observation 6a8aa7f5-c97c-4e5a-bd25-04adefc0fc7c · outbound

This paper cites Long Range Language Modeling via Gated State Spaces.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Long Range Language Modeling via Gated State Spaces

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.555498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:0375b5783688474d013a1bbc2dbfa6f476245b0bd0c620a379fe0379dd257bc3

Observation f6b4b095-2935-4e4d-bbdd-fd1421e6287c · outbound

This paper cites Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Universality of Linear Recurrences Followed by Non-linear Projections: Finite-Width Guarantees and Benefits of Complex Eigenvalues

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.560747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:a84362862752fb036086c799f29c1ade97a95b2e3dcaecb5c92bf96ae63e56d5

Observation 18c879ba-ebab-48c9-b250-077258870c42 · outbound

This paper cites Hyena Hierarchy: Towards Larger Convolutional Language Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Hyena Hierarchy: Towards Larger Convolutional Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.566956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:1e511d69112596ecde3384545faa51466801055c65e780649e5bcac677140e0e

Observation 992899c5-c7d4-46d2-8001-8d4a8435d8cd · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.573112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:ec7da882c70b1bbf6eae6e482862939c9b09ab44ce166233df3165976a259bf7

Observation 8850c0a0-df23-4695-b680-4f56978a3535 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Fast Transformer Decoding: One Write-Head is All You Need

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.578807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:0870acda6c3389db8de680c577780e485d4957eb187d712d5c4c4697c9c54297

Observation 790bdada-5db6-472b-a1e9-83e7d0303380 · outbound

This paper cites GLU Variants Improve Transformer.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models GLU Variants Improve Transformer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.583949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:9f910423a9fa4791bbd9d9d2c91e30208656408ae43a17d17722090d0cd76fd5

Observation 3ff9f499-89a9-46ad-9ed4-a3bdf03cfa41 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.592017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f3e060fead95e411e9de7eb011ec1ec351628862fdd4f2cc5792de5af9fa2634

Observation 73b1b397-82a6-4ed8-9f81-2f94adb8ff41 · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Simplified State Space Layers for Sequence Modeling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:16:11.899530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:ea492d41e961205ce88d0baefbeeea899ab6928b95c1f8cf7ccc314a5d46b050

Observation 275ae357-f483-4b42-a1b2-b4e2a2ec6245 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.603924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:9d6ec225442cdccc821cf7b2d60a82988e40487f99bdb091a516a9c23dc12e96

Observation f3a76fb5-52a8-4a4c-9ab1-75af503ea322 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Retentive Network: A Successor to Transformer for Large Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.609791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:6e793d865440eca6c4628e1c3ebb7b97c122cdaf1fbbaf403a3054da2f442467

Observation 9c1fdd02-67de-4fbd-b50f-4c9d2e9dc6a2 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Long Range Arena: A Benchmark for Efficient Transformers

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.429964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:33d44c40cab022c7435c9a67fe117d5a2738830ed7b0402e503a82f56e59ad3a

Observation baa54c77-d192-4123-9937-305d4535e876 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.435569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:a27cb41b408d579888bbfc4a9a4586dd659caf7088469f5297550767212871c0

Observation eedd0109-dc5b-4176-964b-d1704c146855 · outbound

This paper cites MambaByte: Token-free Selective State Space Model.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models MambaByte: Token-free Selective State Space Model

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.441699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:f68f6ec2e85577b6d2fb71382a07478d9cb9cf00969ab9a1199254f0a366abfc

Observation 61b73720-64f8-4586-9849-e616c6cf35f9 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.450878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:d514a3151fa6ac41677a020d55e7c9bc38d6f368bc275be7dc1ee27e98965f4f

Observation 9073705e-3dfd-4def-a342-4df21a38486d · outbound

This paper cites An Attention Free Transformer.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models An Attention Free Transformer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.456663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:7e1e43b9bfc4a3cbf215d091eccbcf8530d09972a77c67b4b2d8f3c3c9bbe970

Observation 2c6f7f69-0934-4197-bfbc-f9f3fd7320ec · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T06:58:17.463618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:8e010e6d768a38fee19db4520bda59050f65939fd8cd6383a41a3a5610192ac7

Observation 826ca3f4-8586-499a-ab66-85fae5056b78 · outbound

This paper cites (13) We mark all complex variables with˜·for clarity.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models (13) We mark all complex variables with˜·for clarity

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:58:17.634022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:00041d29a8e90e32225ef444770be31980be59400ffb1b1fbdc63c709d64ee74

Observation 69dcc9dd-59ff-4712-a0e5-dec63b133fad · outbound

This paper cites an unresolved cited work.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-15T06:58:17.621278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:21d91a2034c21e89a58bcf81e9be73189c01815a8ae3620f4e553e2e40ec57be

Observation 9cf9927a-fe44-4c45-9151-5c9b99f03a47 · outbound

This paper cites On the left, we compare the performance of different models trained with sequence length 2048, evaluated with a sequence length of up to 32,768.

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models On the left, we compare the performance of different models trained with sequence length 2048, evaluated with a sequence length of up to 32,768

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T06:58:17.625874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:58:17.370396Z digest=sha256:c2f44ad4e1ed2f759c3bd8ae85819508f224f16251292d8f06bb5e39ad47510f

Pith citing papers

Observation 06955ea7-2766-403c-925b-c1127bbe829c · inbound

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality cites this paper.

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T12:16:25.390683Z digest=sha256:7bb28bec42c1fd67a00ebda51d28b26c12da6c8b7503ae67b521a379b17907e1

Observation 2ea19612-848c-4781-bfc9-b911c4afd93d · inbound

An Empirical Study of Mamba-based Language Models cites this paper.

An Empirical Study of Mamba-based Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:31:03.952121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T10:31:03.777169Z digest=sha256:7feb97a63b5d75dfd8909eb07b50518e00e32cd4ac05784196aff3eaf7a48e6c

Observation c3d1962c-1f0b-4826-be5d-5803d408ee51 · inbound

Learning to (Learn at Test Time): RNNs with Expressive Hidden States cites this paper.

Learning to (Learn at Test Time): RNNs with Expressive Hidden States Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T05:20:12.134340Z digest=sha256:a5220338ac1df278aca1785fde909972784d38925bcb757e81cd28269b337510

Observation 33723dc7-91db-4d02-a714-deb19bd93a21 · inbound

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation cites this paper.

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T18:33:19.409017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T18:31:35.391674Z digest=sha256:2289d74a2e7c2773ba5613cf305c889fdc1aa0dabefcb444770b907af3d0a6ec

Observation 00c89e75-9697-4269-b1d7-89f40bde8b1c · inbound

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map cites this paper.

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:28:28.652035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:28:28.652035Z digest=sha256:661d03bfaa1aebf15d5a4d106f07c3e1f957709217866603de5da60a215c6d5a

Observation 4ccf5561-4705-4e92-ba3c-7ceb6021ba38 · inbound

Selective Attention: Enhancing Transformer through Principled Context Control cites this paper.

Selective Attention: Enhancing Transformer through Principled Context Control Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:32.680612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:12:32.680612Z digest=sha256:0f22c3919e01f232b7b801049552aa9cf3d83aec4674ac77bd39bdd4378db5ea

Observation 221fdb52-126f-4b54-bf02-399f2e468f51 · inbound

Hymba: A Hybrid-head Architecture for Small Language Models cites this paper.

Hymba: A Hybrid-head Architecture for Small Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:32.762403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:20:32.762403Z digest=sha256:bb7665a5f2f6e68791d71c2510954715563dce0733a7b9bc9f0cdb63e6bbe9aa

Observation 57bfdbea-7147-42fb-af31-eec34fd06814 · inbound

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning cites this paper.

CLOVER: Cross-Layer Orthogonal Vectors Pruning and Fine-Tuning Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T12:13:55.170833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:13:55.170833Z digest=sha256:43554e097eca386a26acf9ac696472a9902930d2d01daeb1ca0225e5e48fd5e2

Observation c080b9ca-69ca-40fd-978a-c69ae7276dd0 · inbound

Attamba: Attending To Multi-Token States cites this paper.

Attamba: Attending To Multi-Token States Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:56:33.047492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:56:33.047492Z digest=sha256:0b41fd2e5d27ca5a337272c134b4db1ad629834c65b4b74e53e1140fe68cacef

Observation 45fc5245-8876-4b22-b5af-21bb4c1cdacf · inbound

Marconi: Prefix Caching for the Era of Hybrid LLMs cites this paper.

Marconi: Prefix Caching for the Era of Hybrid LLMs Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T10:19:00.101454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:19:00.101454Z digest=sha256:7aecfa51d77c7f9cd912b4d262a5e8db923a9088098e6a8de87cbf34c1062ce0

Observation a6cb877e-6a1a-4b05-8e12-be1c1d90dc4f · inbound

MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance cites this paper.

MAL: Cluster-Masked and Multi-Task Pretraining for Enhanced xLSTM Vision Performance Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:45:21.875332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:45:21.875332Z digest=sha256:410ee2855cdbe5c8ec440d622cb606ca0d56392db795c938c279ad8679dbf63a

Observation 35461595-eefe-4b9b-b199-58a8ab6265e1 · inbound

Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models cites this paper.

Expansion Span: Combining Fading Memory and Retrieval in Hybrid State Space Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:17:49.920570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:17:49.920570Z digest=sha256:289da5c0d5ee3f53128832c925abcc20f977f77722d5c2ff44ddc95f1a43848a

Observation 68e731b5-6662-4f69-862a-7e80fd990183 · inbound

On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages cites this paper.

On the Expressiveness and Length Generalization of Selective State-Space Models on Regular Languages Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:46:47.564683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:46:47.564683Z digest=sha256:095e9dcbcec8787b5f605f79c0334aacd17c0ebf2ce42e5d141209400ac43163

Observation ff742bf8-2803-48b4-990b-dd784d4ac900 · inbound

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing cites this paper.

Understanding and Mitigating Bottlenecks of State Space Models through the Lens of Recency and Over-smoothing Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:33.407687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:33.407687Z digest=sha256:672056631ac64d8fb67faa07afc81160b3c4f323d4706d91be5ab0b40d01bc39

Observation 4ad863bf-577d-4590-9214-3901ded4aafd · inbound

Titans: Learning to Memorize at Test Time cites this paper.

Titans: Learning to Memorize at Test Time Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:08:14.982302Z digest=sha256:e3616803d34a634705869798da3ab43ce41fc485df35dd8ddfc0f445e9ddb699

Observation 44ea67d5-8c86-4734-b520-6a1dbc847437 · inbound

MSWA: Refining Local Attention with Multi-ScaleWindow Attention cites this paper.

MSWA: Refining Local Attention with Multi-ScaleWindow Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:39:31.411825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:39:31.411825Z digest=sha256:2f5d2c312826634a0ebb1f87baac9cbc794d32aa31c81c9ff7cc263b73e06475

Observation 7ba4b453-34a9-4e90-a3e4-bb2cb35c86f5 · inbound

Test-time regression: a unifying framework for designing sequence models with associative memory cites this paper.

Test-time regression: a unifying framework for designing sequence models with associative memory Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:07.027580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:22:07.027580Z digest=sha256:275b98c3922a9fe8d93f2c3c94a44f4091992d6263f84e4e0a2360f4e7f8344f

Observation 64bd7a44-1de4-491a-963a-53bee77dd574 · inbound

GRAMA: Adaptive Graph Autoregressive Moving Average Models cites this paper.

GRAMA: Adaptive Graph Autoregressive Moving Average Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T16:57:29.078058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:57:29.078058Z digest=sha256:460a8f6215de2303f00889aa70a742888649f3035d57f75f61325659e9438eff

Observation e9a88c47-c086-40f5-8f71-0df4b92ab72e · inbound

Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing cites this paper.

Explore Activation Sparsity in Recurrent LLMs for Energy-Efficient Neuromorphic Computing Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T21:19:23.945339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:19:23.945339Z digest=sha256:46048f460192f69bffb6f82e4190725e873d984854f680ea119f005b87b40878

Observation 056e3391-19a9-482c-8f81-e5c0ee415e9c · inbound

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach cites this paper.

On the Expressivity of Selective State-Space Layers: A Multivariate Polynomial Approach Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T13:05:31.591911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:05:31.591911Z digest=sha256:3886652d74f251ed266410d362533e598cccf51f211a025e5a18444ed7ed620d

Observation b16394b4-5d9a-4a44-8eae-af6350b2a513 · inbound

An Uncertainty Principle for Linear Recurrent Neural Networks cites this paper.

An Uncertainty Principle for Linear Recurrent Neural Networks Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T22:10:15.888345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:10:15.888345Z digest=sha256:d14175f903f33ee8892f3addbf282fedf0dd0c074ece8983886855d9bc63cb51

Observation be4a3871-60c1-495e-ab58-550edd005235 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 58

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:46:30.073118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:1675c86481689d3ddf711b6a06e7d520014e541f2f7bfdd9cc30007984fde20a

Observation 0320d96b-4f2a-41e7-99e8-7e1a7cace126 · inbound

It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization cites this paper.

It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:30.286857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:30.286857Z digest=sha256:dfea3c105e086b6b7434ec068c018b53b6f3461c28d25d951dda21902401d837

Observation e208c05f-fcd9-4ad1-babd-562f36e5412e · inbound

ForgeBench: A Machine Learning Benchmark Suite and Auto-Generation Framework for Next-Generation HLS Tools cites this paper.

ForgeBench: A Machine Learning Benchmark Suite and Auto-Generation Framework for Next-Generation HLS Tools Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:34:26.698505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:34:26.698505Z digest=sha256:408ce83ce95457a73bafe5d54a5b516b4bf4ff3ee6e74acd8253f65e152948f3

Observation f2dfb2a3-992d-46be-9a70-efcac9a03df4 · inbound

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities cites this paper.

LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:16:29.630423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:16:29.630423Z digest=sha256:2770f77db080311c85616f1ebdea72bd3aa017645f784a3bffcc66a19ddb6c9a

Observation 0fa4d082-08dd-4cc4-90f1-f21da8c1df1d · inbound

Quantifying Memory Utilization with Effective State-Size cites this paper.

Quantifying Memory Utilization with Effective State-Size Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:22.403929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:22.403929Z digest=sha256:d7000775337dea0ac5ca5bc4276b534167891366d548a1280dad6a3ddabc4149

Observation 69fb5a92-dc91-4efb-abf8-271cf9a58e50 · inbound

Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook cites this paper.

Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:07.702962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:07.702962Z digest=sha256:1e7efec0d21e5998863af388e8ffa528c38c22891e005b35d28b4a881bb8f02c

Observation def813ee-7313-44f2-919a-5f6881eec9ed · inbound

Reasoning Capabilities and Invariability of Large Language Models cites this paper.

Reasoning Capabilities and Invariability of Large Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T04:40:57.775150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:40:57.775150Z digest=sha256:1279423b3ddddc05431c382e4fc241d33bf659850f293629222a83c55ba5cfea

Observation ced32ed0-de43-45dd-84c4-b0d6b4dff9f9 · inbound

Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling cites this paper.

Message-Passing State-Space Models: Improving Graph Learning with Modern Sequence Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:36.596300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:32:36.596300Z digest=sha256:1db94b818011b84556c2a8a1108d1d75654673b814ce62a60a3037c9cbb568d8

Observation 6e69752c-4362-44d3-abba-04e2c13ff613 · inbound

Revisiting Glorot Initialization for Long-Range Linear Recurrences cites this paper.

Revisiting Glorot Initialization for Long-Range Linear Recurrences Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:08.747322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:08.747322Z digest=sha256:0f5052ac9c6d808836430b13ff5b465eddc6f4cc78652caada172043ef354f56

Observation 4e84e033-a2a6-4b5f-af98-82f51edd9fb1 · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.020020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.020020Z digest=sha256:dd37991c1a1d52ed0abd5fc8416ec0cd5230ed7d4bc9a1cec26847af2cebed3b

Observation 62752f89-cdd0-4fdc-a32d-e0684ef8d378 · inbound

Geometric Hyena Networks for Large-scale Equivariant Learning cites this paper.

Geometric Hyena Networks for Large-scale Equivariant Learning Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:11:41.546011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:11:41.546011Z digest=sha256:8df9e820edb6ed023e5711f626d2eacbae01fcb2cd8bbe1b0dbb42d293ae6c7d

Observation 51c5621c-26de-4fc5-a3fb-9b8fc8667510 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:12.746450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:12.746450Z digest=sha256:dd217871d73caf71d762fb9610130827afa5d613e6fc9be3ad85208d3f00753f

Observation 55da58e0-17ae-4336-a8a0-252aad8ba740 · inbound

A Survey of Retentive Network cites this paper.

A Survey of Retentive Network Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:56:23.523266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:56:23.523266Z digest=sha256:8afba811419e6b4b482d171985691e521fb631c3361322cc32e379ac3fad6bea

Observation e10fe4ab-9031-49e5-a69c-d6f1a50e3800 · inbound

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection cites this paper.

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:44.335630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:44.335630Z digest=sha256:f222015add5318992b45d5dd12a1387055ac34bb1dd62c36999cae5c1f10e9d2

Observation ebc456d3-14e0-4f9f-ab24-06bafc44b2d6 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:17:24.628777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:17:24.406028Z digest=sha256:208368b53ea1d49dfdc9fcc06f21a1f23769d631084390c939991d923b6caa5c

Observation c333a267-75e7-4be8-bc80-e3cfba4385c1 · inbound

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent cites this paper.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:40:11.917229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:40:11.917229Z digest=sha256:afd1c0624f75241062833644387adcacd0a5e5fbedf62d18e815db5b72a5fc8a

Observation 889cdfca-77ea-444e-be11-5f1d974dba96 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.624319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.624319Z digest=sha256:38b079d4ceddd1d75a28148bd413273337f8f5ebca4caec3fd66712a910d0898

Observation 3434b521-120a-42f8-ba19-6eb9e017bf1d · inbound

A Systematic Analysis of Hybrid Linear Attention cites this paper.

A Systematic Analysis of Hybrid Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:09:55.788354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:09:55.788354Z digest=sha256:dbac7387fecffe24a69911723d073b67be05da316de0589081b95301d124810c

Observation ad621e97-2937-4b58-bf06-027b3959d7c1 · inbound

Lizard: An Efficient Linearization Framework for Large Language Models cites this paper.

Lizard: An Efficient Linearization Framework for Large Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T04:42:04.534710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T04:37:55.034479Z digest=sha256:2166c8d9713a4868a626d454a07ab26a25ef5058082582eee8b827bc931df922

Observation 1391d620-ee6f-4a0c-9b19-b0c38e309ee1 · inbound

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance cites this paper.

Falcon-H1: A Family of Hybrid-Head Language Models Redefining Efficiency and Performance Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:44:04.925289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:44:04.925289Z digest=sha256:41b56247e1b79909d03d7630918e6ba8cf8f7a8c275e450392bb4b0a429d909f

Observation 3708a5fe-043a-4854-917d-3b1719bd4270 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-18T18:51:45.775697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:6bcc45e4647d4e63f05762ab30dee3a4367ddc0f63274fa40078dea6f70d8237

Observation 9f1ea1b3-5a5c-4b25-ba5a-f1778b03a762 · inbound

Elucidating the Design Space of Decay in Linear Attention cites this paper.

Elucidating the Design Space of Decay in Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:29:21.297289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:29:21.297289Z digest=sha256:01aa139f80847fb20206f8d99cae2eba202d80215dc4a95b089653f7d5a26c7c

Observation 48db92f2-b4c7-4b96-867b-744794c02828 · inbound

Short window attention enables long-term memorization cites this paper.

Short window attention enables long-term memorization Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T12:11:21.808263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T12:10:42.646127Z digest=sha256:cc6fec856536860ac2ae479d1cd7b26353a3f3225cff164d2a62da7ba5cab04c

Observation 82a3cfd6-46e4-4b09-877f-f83ccafb0747 · inbound

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights cites this paper.

Hybrid Architectures for Language Models: Systematic Analysis and Design Insights Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:21:15.108907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T10:18:04.431436Z digest=sha256:c62add0d61af9c308f3fab002d71ef2d8899718b536f776dce446a31b942021b

Observation 60ec5ed0-64e6-4810-b69a-d0a5625e074e · inbound

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone cites this paper.

DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T21:20:34.467266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:20:34.467266Z digest=sha256:3f81100b6427d22dcdb68b528954847677cc859eda1f2bf63a3d76cba0fbb19d

Observation 863a79ef-9659-4fb0-b88e-684dc47799f3 · inbound

Selective Rotary Position Embedding cites this paper.

Selective Rotary Position Embedding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:40:14.861957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-17T20:36:49.650895Z digest=sha256:aa52c39d43c1ae3cd146537f3a9fc33993559108b32c2b6dfea863f33364960e

Observation 499c3aa2-5ce8-4b8a-b40f-535f35f94d3b · inbound

Selective Rotary Position Embedding cites this paper.

Selective Rotary Position Embedding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T21:03:18.981806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T21:03:18.981806Z digest=sha256:71551391c7aef5895e27b9b3dffa3822eb28d6b230eebed05d85f0187b921190

Observation af9b1816-3b38-4915-8786-0b4661921526 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:00:27.115409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:3ef0443df22e1fc3dd691742db2171e19e0f6fe6e8f416f1774e39345316d64a

Observation 6f226ae0-8f0e-4025-88d8-a9f70e0c8272 · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:48.561819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:48.561819Z digest=sha256:63989ea583b646c274940131622e0f5859570cc631773cc5837175b0095b0a89

Observation 3b83e026-48db-4285-9fa0-0126f68688f8 · inbound

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction cites this paper.

Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T10:12:43.396788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:12:43.396788Z digest=sha256:accc1f5a94d3c80e2abf06a03b9cfb236d3e2980c19fd5a32a3c6a1272ccf6b4

Observation 18a0e62c-33ee-474a-9528-d0cb10165e65 · inbound

When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models cites this paper.

When to Think Fast and Slow? AMOR: Adaptive Entropy Gate for Hybrid Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T11:50:52.802578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T11:48:50.587732Z digest=sha256:77e9ea13100414e1af89748c7595be8e8f524430df58c1dedc59f813d1c6e819

Observation 38ac6792-76ff-4180-82a3-fead463c9dcf · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T21:00:17.897130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T20:59:33.902420Z digest=sha256:2b2bc10eaaae2e2d7904a4b0ac9c75340910800c0121282a2d247f5cfadd8004

Observation d4c25bb4-d24c-44a4-a2f9-991ef4125905 · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T12:50:09.519771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T12:45:27.150368Z digest=sha256:34686ba39c7153ad5399b724a92ba0fc24435c63532685fa4574fa6ff7582327

Observation 98aafb65-e4c0-4f2a-9a27-cbbbbc0b0d3d · inbound

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cites this paper.

RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T22:05:42.987805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:05:42.987805Z digest=sha256:fecc61e6884de6fd43a47e53b2cc5bd94ca3d363533f68b5aabb98884ab32886

Observation daa7580c-ff19-4d6f-8931-76acb0864547 · inbound

When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models cites this paper.

When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:24:41.108227Z digest=sha256:18bf9ab9aab7cc0dd3a96366b59a9329b9694cf63cbc2a1a0623fe7ef58ddf9f

Observation 50eab17c-36ae-4490-acd5-02b0dccfdb7d · inbound

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models cites this paper.

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T17:21:21.451268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:21:21.451268Z digest=sha256:35a35b5b25db42ce85e6fa8e1bb6c7be768af8685c62cdcdfd64b42788f80325

Observation 92a3e1eb-c599-41c1-a96a-74bd0879067a · inbound

LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling cites this paper.

LPC-SM: Local Predictive Coding and Sparse Memory for Long-Context Language Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:25:31.080573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T11:22:08.937342Z digest=sha256:901358cb46dd27758326957d90cb1fce0882a84c14e2c1355ce8982d49037ae6

Observation e0b7d0a9-90b9-4854-b824-0dab5b38ea57 · inbound

Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models cites this paper.

Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T16:55:36.438348Z digest=sha256:cb850ee8a6772cf9cf591744a0a3febdf6967bd101b901afa35be4f0ee4c6404

Observation df3bb33e-85ce-467e-a284-af7124e56847 · inbound

Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models cites this paper.

Mambalaya: Einsum-Based Fusion Optimizations on State-Space Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T16:51:07.301614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:51:07.301614Z digest=sha256:858c86206c599c44f81541cf41d83f6a5e15d33c0143ff5c0c093c38be5db1a3

Observation 9cdce8f7-b0c3-485f-ade1-7cdee9c524c2 · inbound

CAWN: Continuous Acoustic Wave Networks for Autoregressive Language Modeling cites this paper.

CAWN: Continuous Acoustic Wave Networks for Autoregressive Language Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T10:31:11.574398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T10:31:11.574398Z digest=sha256:ba91f79bb640e419c15874e6a33a840331631c7ff50a8442993e5d5d6f7b63d8

Observation 6400ba6b-3a57-4e36-9564-fd52ac10af70 · inbound

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space cites this paper.

Phase-Associative Memory: Sequence Modeling in Complex Hilbert Space Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:39:50.311059Z digest=sha256:cd09829df1b9a50af9d791e414071b492e329c37ee6e4c930541194151d9d6d9

Observation 701eff04-23f3-44ac-b058-ca11003e9a80 · inbound

Optimal Decay Spectra for Linear Recurrences cites this paper.

Optimal Decay Spectra for Linear Recurrences Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T17:21:15.178258Z digest=sha256:4c918650be6d7a87f0b592bf93646695142a085556e8057798e437a3aa7c9225

Observation a32ec29c-a010-4964-94b3-eecb39f70070 · inbound

TAPNext++: What's Next for Tracking Any Point (TAP)? cites this paper.

TAPNext++: What's Next for Tracking Any Point (TAP)? Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:29:15.631211Z digest=sha256:6222bae30c4e80fc51d3e16958dd84c4b7176daf0b2340f8f50fc6eed4316b04

Observation a45c8722-7873-4785-876a-8a40cd86a7dd · inbound

On the Expressive Power and Limitations of Multi-Layer SSMs cites this paper.

On the Expressive Power and Limitations of Multi-Layer SSMs Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T12:36:09.265655Z digest=sha256:1762da4530eb25331f0546cd2d4ecd1d0d3cc61d242a5a9bc5a28f64209ed710

Observation 56d86a29-2489-4db4-b253-f5d44635e2d1 · inbound

Scalable Memristive-Friendly Reservoir Computing for Time Series Classification cites this paper.

Scalable Memristive-Friendly Reservoir Computing for Time Series Classification Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T01:10:45.859345Z digest=sha256:d726216c3ad19acd74012d1ae39826a48c31e7e364db007087d73ede0b41ef05

Observation 75eff084-6ebd-43c0-856c-502724664f30 · inbound

HubRouter: A Pluggable Sub-Quadratic Routing Primitive for Hybrid Sequence Models cites this paper.

HubRouter: A Pluggable Sub-Quadratic Routing Primitive for Hybrid Sequence Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:21:07.816749Z digest=sha256:e761992ec28e5d247337273dd6f5929c262697c3ed168b309c514c2961d9a381

Observation ddd9849a-3219-4fcf-b569-e3919cb12db7 · inbound

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference cites this paper.

SpikingBrain2.0: Brain-Inspired Foundation Models for Efficient Long-Context and Cross-Platform Inference Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T12:18:23.898779Z digest=sha256:f085e9753bb7225e9a95db31895de6c349f69a7382117444bcda343fc5e3154d

Observation 089576dd-8998-4e0e-a175-7547cbf926bf · inbound

The Impossibility Triangle of Long-Context Modeling cites this paper.

The Impossibility Triangle of Long-Context Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:17:44.033300Z digest=sha256:19a6ab16ebc0fd4bb50a266347d994ac54f5c242aeb6ef9c806f557a4dce6a94

Observation 7340369a-81bb-4691-9c29-ce1f099bc359 · inbound

How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences cites this paper.

How Long Does Infinite Width Last? Signal Propagation in Long-Range Linear Recurrences Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:52:21.272304Z digest=sha256:e0da46195437747c092c5fbfac4a4d54ea5b0536d168932d209c1db6cba3b23e

Observation 4377f820-c8ca-440a-83e0-675a877b3276 · inbound

A Robust Foundation Model for Conservation Laws: Injecting Context into Flux Neural Operators via Recurrent Vision Transformers cites this paper.

A Robust Foundation Model for Conservation Laws: Injecting Context into Flux Neural Operators via Recurrent Vision Transformers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T16:52:57.496945Z digest=sha256:b90baba78447953669640fe637522597b63a29ca164f7148a7b8783ab3b95b1d

Observation 64cc7df7-1c6f-4b56-b12a-afdc3eb9d202 · inbound

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention cites this paper.

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T15:27:55.566795Z digest=sha256:1abe079fac17eab8f15ed18d0f3f60f1c34ea93e9db1bf245d222ad9ae1aaf99

Observation 96e8e322-c921-43a0-b314-a6df5715f1a1 · inbound

Priming: Hybrid State Space Models From Pre-trained Transformers cites this paper.

Priming: Hybrid State Space Models From Pre-trained Transformers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c3f8107be0adf44bc4bf2549edd9a432c02f8a0f9ffc1a8a18a8fbf12180325d

Observation 790946f8-6032-4a68-b1ba-7269b092e9e3 · inbound

Kaczmarz Linear Attention cites this paper.

Kaczmarz Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T01:15:58.330766Z digest=sha256:75ae542a00acf22061bdf90bfa4e218f6b4e7b83c2b15931431094a279ab747a

Observation 401870ed-087a-44c5-b4b1-d6c4e4c9aeef · inbound

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading cites this paper.

MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereading Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T05:22:43.330891Z digest=sha256:9db82c20f12adddd0ee82ba7fc2a1d1c2d679e8bc098e9575ec0bd7a3d232600

Observation 9dfb6fda-8d60-4f2c-b257-bcb8a0fa20ed · inbound

A Single-Layer Model Can Do Language Modeling cites this paper.

A Single-Layer Model Can Do Language Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:40:12.906234Z digest=sha256:42d01f660bd12948b01dab64c19f75b6647ef27ee50e582a6d6175834cb1c735

Observation 83d8a01a-56d0-4642-b16b-1c7036365860 · inbound

HexagonalWarriorMamba: Superior Threshold-Dependent Multi-label Classification of 12-Lead ECG Cardiac Abnormalities cites this paper.

HexagonalWarriorMamba: Superior Threshold-Dependent Multi-label Classification of 12-Lead ECG Cardiac Abnormalities Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:14.929076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:55:41.887085Z digest=sha256:e02466fe63153cbe6ea1ea779c5003cd2f932dc9451c053a40c6f56a282e756d

Observation e1fae2d7-c4ea-4e00-96f4-2c0b0b29b0dd · inbound

The Routing and Filtering Structure of Attention cites this paper.

The Routing and Filtering Structure of Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T22:09:07.720892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T22:04:15.518405Z digest=sha256:b9aa8c4f5990dda8296d56a0325e4ef6e58ca429ef38fd2c1fe9a93384d04cb7

Observation 0c87302f-66bf-4d22-bba4-ee4092b054db · inbound

Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models cites this paper.

Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:53:15.168578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T11:48:42.834602Z digest=sha256:dc0e4e35c27a9928f30515023a25f98506fb35035bdb15de914abfad6afe2dd6

Observation 4aa09321-2184-4788-ac6f-37eaa4c3edd2 · inbound

Towards Understanding Self-Pretraining for Sequence Classification cites this paper.

Towards Understanding Self-Pretraining for Sequence Classification Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 101

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T05:33:58.868560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-21T05:29:58.809024Z digest=sha256:a7bf1c5b98eae1c35d22d464bcae93521680d0ad505734bb561feae58992d896

Observation e4cfe493-1848-4837-9464-cef8867d4bcf · inbound

Interdomain Attention: Beyond Token-Level Key-Value Memory cites this paper.

Interdomain Attention: Beyond Token-Level Key-Value Memory Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T14:04:44.092256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:dc725edd07c5b8d0d9a21ea7283d6affa776076876e0773be38b591130543663

Observation 9fab381a-a2bf-4a65-9094-474dbe26edb4 · inbound

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference cites this paper.

Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:43:59.493853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T21:37:51.638904Z digest=sha256:635cea77e418a4a7e54a8596c6af9a7acfde263764fa55ecc754a8bce98a32d8

Observation 2efbb355-3685-403c-8cb5-003ce1c758cb · inbound

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior cites this paper.

Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T19:43:54.937865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T19:36:13.559393Z digest=sha256:8c31f433146144fb0881f97b16cee3b2d5edfe3e2ce466d998d6755a45f97243

Observation e3f096b0-1cbd-4357-b8e1-626a9d28993d · inbound

Memory by Design: Probabilistic Sequence Layers cites this paper.

Memory by Design: Probabilistic Sequence Layers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T20:26:12.920120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T21:07:31.407554Z digest=sha256:a5c0ec61438c394c723a951cbbd60e2e3f4283dde789371b0cdcc3ac00405a4f

Observation 48ca4380-ab64-4ea9-aab1-9d50ceb66ddb · inbound

Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing cites this paper.

Trading Complexity for Expressivity Through Structured Generalized Linear Token Mixing Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T23:02:45.801805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:e332485c92470fc3ba6fde7d865a505beacf0d691dd8f21fb72c3ad9802da6dc

Observation 88cd020b-e37e-4376-b097-65ad34fab426 · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:21.029340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:eb24ead9880aa63e4019205665f61540339c37779be2affcf2719d6f5a4ba122

Observation 02f27307-df19-4fa0-8a6a-6ce6cce7549b · inbound

LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling cites this paper.

LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T06:36:44.182382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T07:19:31.075298Z digest=sha256:97e565af8fd7848137929fc19b93a8f3711cb352b9cf8a815df9c25dff57ae86

Observation 324b4bc8-0224-4646-9c55-f236ee2be205 · inbound

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition cites this paper.

Titans-as-a-Layer: Test-Time Memory for Conversational Speech Emotion Recognition Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T23:07:26.839436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:28:44.652813Z digest=sha256:72657c8459e1e02513f05de0a242d025d102f9de2c24632f89ab3d67111b8c47

Observation b25d24ce-a02b-48a6-b3bd-039a3bc986c2 · inbound

Blurry Window Attention cites this paper.

Blurry Window Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:46:14.017951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:43:34.429061Z digest=sha256:bda2f85736cc7018179e6cdefcf965f9b234a299454f98d2ba3dd44bceb7d999

Observation a25b4ef6-89f8-44d7-b979-031b11ffa75b · inbound

Free Parametrization of L_2-Bounded Structured State-Space Controllers for Nonlinear Control with Stability Guarantees cites this paper.

Free Parametrization of L_2-Bounded Structured State-Space Controllers for Nonlinear Control with Stability Guarantees Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-27T12:10:53.818017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T12:08:57.506333Z digest=sha256:7bbeac18529268067e1a4d3822fcf011efaebaf34a4f89b56f0102a63820a036

Observation e1a196cd-9290-4ccc-96b3-02e2ec849791 · inbound

PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation cites this paper.

PLUME: Probabilistic Latent Unified World Modeling and Parameter Estimation for Multi-Finger Manipulation Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:17:41.379799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T12:50:32.341334Z digest=sha256:d8f3d02cfada6b2b3328f309d41e8d05ca9dcf0be1ac13cf33c970543dc0eadc

Observation 63a05911-a45c-4fd6-a0f0-b66700b73686 · inbound

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning cites this paper.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T09:47:59.882874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T10:18:54.163862Z digest=sha256:73e4e343f2a8d022ddc8786765344b5c65d82835f24642446e3474244f73de6a

Observation cef66bdb-d22a-40e8-a56b-ae62af49193e · inbound

Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution cites this paper.

Linear Recurrent Unit with Semantic Modulation for Image Super-Resolution Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:29:29.435897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T18:06:22.521551Z digest=sha256:7f0f25dd6ad226296aff12493db0ec7cfc3ad7b5f9024e33e17084a3d8de8ce3

Observation 56711364-b5c6-438f-b08b-655bbb5d52a3 · inbound

Tapered Language Models cites this paper.

Tapered Language Models Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T09:59:45.305176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T09:11:20.341634Z digest=sha256:32af56a4b6071447191580ea83276b51205f8b1ecd8409fd456b0a8a85a15224

Observation df893389-7fa6-4dbe-bedc-a57a55a99500 · inbound

Harmonic: Hierarchical State Space Models for Efficient Long-Context Language Modeling cites this paper.

Harmonic: Hierarchical State Space Models for Efficient Long-Context Language Modeling Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:42:35.933085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T18:55:19.253621Z digest=sha256:29370b1a2380af0ccbe5be061eeb6fc54a33a33eda89e0e1c342231da4221dbd

Observation fdc4b206-788c-4312-b5af-077c67d64a90 · inbound

SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning cites this paper.

SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T15:19:56.025588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T01:46:44.408246Z digest=sha256:19a7b13b27346a31d466b8825608d43bf76d18de7cc8c0efbe28da93ac500470

Observation c128b580-bcc7-4126-b36b-bf9276309c66 · inbound

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention cites this paper.

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T14:09:53.441851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T04:26:00.066003Z digest=sha256:9633cd66997b1c42b4c56639219367852c8ae7756018442344312776c773d4cc

Observation 2f81bd82-e038-42cd-a18e-3791e2cff5fd · inbound

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention cites this paper.

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:44:37.449959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T09:42:52.096671Z digest=sha256:04800a1c2e8b9aa7c4c3b450ef5df5c5f52f2a1cfba86f3085bc00b22e650af5

Observation 21955ff2-0e38-425b-9bf5-0f884c42ab49 · inbound

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention cites this paper.

CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T11:49:31.618283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:49:31.618283Z digest=sha256:2f476c5eb984ae465eaa91b427dd8c98ccf0cb3edb61ac99985e151fab8c7910

Observation 4c34e4da-8684-491b-b038-bd17202c16f6 · inbound

The Context-Ready Transformer cites this paper.

The Context-Ready Transformer Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T18:25:58.296712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-29T01:59:43.093422Z digest=sha256:84f56f41c9f3c97098d2558cb5965fb4f99be32123ea15e70049b910e88d7323