Pith. sign in

Paper Citation Record · LEDGER

Priming: Hybrid State Space Models From Pre-trained Transformers

As of 4 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 3 inbound Pith citation observations for arXiv:2605.08301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.08301 v1

Coverage vector

measured 100 of 103 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:14:01.584159Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T02:03:22.318771Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 103 outbound references displayed

  • verified exact31
  • verified fuzzy58
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ff7aafc-a2be-45bc-bef2-5cb6822ee13b · outbound

This paper cites GQA : Training generalized multi-query transformer models from multi-head checkpoints.

Priming: Hybrid State Space Models From Pre-trained Transformers GQA : Training generalized multi-query transformer models from multi-head checkpoints

Reference 1

Resolution
verified exact
doi, observed 2026-05-12T01:16:13.883063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c9c0d0d3580ced57c2da9ab18fee7f37ccc26b17e56d88b7a414fda762d9f162

Observation 51e8d96c-ee7e-4d16-a986-214fc6daf6c7 · outbound

This paper cites Training-free long-context scaling of large language models.

Priming: Hybrid State Space Models From Pre-trained Transformers Training-free long-context scaling of large language models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.628756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2e18c1fbae3b984307b479a11aba8ad5186b79910cb1267c04f74646bedcd957

Observation b5a230cb-4f44-423f-bf26-dde20eb946b5 · outbound

This paper cites Zoology: Measuring and improving recall in efficient language models.

Priming: Hybrid State Space Models From Pre-trained Transformers Zoology: Measuring and improving recall in efficient language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.623082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:48f99a936e9432ac27c37f1026f6ea15b478774504eb173735ee2c4a3aa519bf

Observation d32423a0-737b-4ff9-9aa2-8ef27a780df7 · outbound

This paper cites Hybrid Architectures for Language Models: Systematic Analysis and Design Insights.

Priming: Hybrid State Space Models From Pre-trained Transformers Hybrid Architectures for Language Models: Systematic Analysis and Design Insights

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.319329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:774954c5675ef662bd4faa37a7fe8fcf13d558018461ee9f7e593fc4946ca8b6

Observation 937d8321-0e24-474d-b17f-f5c9de20b1f2 · outbound

This paper cites Chan, James Demmel, June Donato, Jack Dongarra, Victor Eijkhout, Roldan Pozo, Charles Romine, and Henk van der Vorst.

Priming: Hybrid State Space Models From Pre-trained Transformers Chan, James Demmel, June Donato, Jack Dongarra, Victor Eijkhout, Roldan Pozo, Charles Romine, and Henk van der Vorst

Reference 5

Resolution
verified exact
doi, observed 2026-05-12T01:16:13.855279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:12d2a3cbe98f033c5b1f63b5e5efb4b738b0181e51017e049d01dff8350a39f4

Observation 6f681267-1086-4cd4-bc16-65167b3d0426 · outbound

This paper cites NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model.

Priming: Hybrid State Space Models From Pre-trained Transformers NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:02:08.433616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:eea981bc45babf6702785470a61b52579401d7b57d6cff84de18377e020d205e

Observation bfc6085b-37c1-430f-a58a-b651690c2097 · outbound

This paper cites o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \.

Priming: Hybrid State Space Models From Pre-trained Transformers o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.624776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:1a6da9db135d074300fa6bf5f08db8c3e9e366b6a80a70a5f6cc1105fb0c9619

Observation 0cb57d44-57fe-44b8-bfb0-692de7f97051 · outbound

This paper cites Transformers to ssms: Distilling quadratic knowledge to subquadratic models.

Priming: Hybrid State Space Models From Pre-trained Transformers Transformers to ssms: Distilling quadratic knowledge to subquadratic models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.619608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d8771be49e7a086bdfc9775b1a03b83e3ffbf7c5d60c3ab33062fd8ed3e0f584

Observation 81eea1c2-39bc-401c-bff5-fbec4135112c · outbound

This paper cites NVIDIA Nemotron 3: Efficient and Open Intelligence.

Priming: Hybrid State Space Models From Pre-trained Transformers NVIDIA Nemotron 3: Efficient and Open Intelligence

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:40:43.189893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:301d033938e7c12f0ac920bcfbe3eddadf9aca9cf5c1a64ec33da9a3954ca72b

Observation b9810ab0-561f-44cc-be02-b1746864dd73 · outbound

This paper cites Nemotron 3 nano: Open, efficient mixture-of- experts hybrid mamba-transformer model for agentic reasoning.

Priming: Hybrid State Space Models From Pre-trained Transformers Nemotron 3 nano: Open, efficient mixture-of- experts hybrid mamba-transformer model for agentic reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:24.383814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:eb52c063bc36bc28551cd90b8d378868d6cb09fdda1b00794497894b62e8322d

Observation 5d2e8985-7c82-48f3-9dfd-757aa98d9d5e · outbound

This paper cites Qwen3-Coder-Next Technical Report.

Priming: Hybrid State Space Models From Pre-trained Transformers Qwen3-Coder-Next Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:12:48.468873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:01fc875a9fb4d2a89bbee31a39e6c6d23bdc28f54f8aed3df78bb2880bf8d836

Observation 73390f12-bb59-4489-a813-743d78cc460c · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Priming: Hybrid State Space Models From Pre-trained Transformers MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:4fab0c24a1de6a2db4a7878d3f31c488fcb0cebaea765dcf7c3144537d68e51f

Observation d60bf531-4344-49c9-921a-7ce4e222f91f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Priming: Hybrid State Space Models From Pre-trained Transformers Evaluating Large Language Models Trained on Code

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.324208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:aa815f7d51efeb12ba3026f0b0233af3dbb1a5fa87e945298d2abe8c286ab769

Observation 017c088b-7c25-41f6-81e3-d757de190138 · outbound

This paper cites Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun.

Priming: Hybrid State Space Models From Pre-trained Transformers Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.617611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c922b5e36854242dacec3001bb29f655ffb29a747219b6db8ce0180d0ba5457b

Observation 3020041d-d596-407a-8758-7b23be2f8263 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Priming: Hybrid State Space Models From Pre-trained Transformers Training Verifiers to Solve Math Word Problems

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.428564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:4ec51f28da4acdbdf278d2ef2214ba2bfe5082077ce0924d4999c7176549c7d2

Observation 7e8955e8-2cc8-4860-8c1b-c175cfb36fbe · outbound

This paper cites Albert, Pranesh Srinivasan, Haining Pan, Philippe Faist, Brian A Rohr, Michael J.

Priming: Hybrid State Space Models From Pre-trained Transformers Albert, Pranesh Srinivasan, Haining Pan, Philippe Faist, Brian A Rohr, Michael J

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.621362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:eefc125f08ea3f9405ba6af61422281287de8aed2fd8186e11b641ea3184ba10

Observation a78944fe-cdae-49a4-9fc0-fb2af640fb0c · outbound

This paper cites Transformer-xl: Attentive language models beyond a fixed-length context.

Priming: Hybrid State Space Models From Pre-trained Transformers Transformer-xl: Attentive language models beyond a fixed-length context

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.626741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:ad244b5c3d5ed13e5c41d1ec1c2da99bc0535facee780d89bdac3c0872e0fb2c

Observation bd33fdd7-d439-45d6-a3c7-359a5313b4ca · outbound

This paper cites Transformers are ssms: Generalized models and efficient algorithms through structured state space duality.

Priming: Hybrid State Space Models From Pre-trained Transformers Transformers are ssms: Generalized models and efficient algorithms through structured state space duality

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.630680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:52cee3ff4a39cfc9a452cc8838d60e346f5c7fb0ce0ec89c6508f490206f9234

Observation 01c2b27f-cbdc-4fdc-bcb6-ca9670a9df3c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Priming: Hybrid State Space Models From Pre-trained Transformers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.610375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:48a6ceb2273f31eac59d98fd09f990fb117ce41d0e2db209afc501cac5b9d734

Observation 96e8e322-c921-43a0-b314-a6df5715f1a1 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

Priming: Hybrid State Space Models From Pre-trained Transformers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:58:17.635264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:64a6848321103b8993886b9d158385b5b4e0d2a6051b29104c2a8781ea85f67d

Observation b3df9085-4f0a-4da3-9d68-d551fe8ab05c · outbound

This paper cites Fewer truncations improve language modeling.

Priming: Hybrid State Space Models From Pre-trained Transformers Fewer truncations improve language modeling

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.608426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:560737f503ad2d9d511e2fbb3741a9c09f2db1a18f6d837fb46664326ca78209

Observation 363e801b-12a3-4f15-b0e3-bfc658e1c6b7 · outbound

This paper cites Hymba: A hybrid-head architecture for small language models.

Priming: Hybrid State Space Models From Pre-trained Transformers Hymba: A hybrid-head architecture for small language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.612412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:59c9c9f82f9ff597ef3f769704ced6691db4512eab7e9f6285629aa88df6f704

Observation 279ab5d2-cf54-4277-b35d-d2a7da0b70c5 · outbound

This paper cites A mathematical framework for transformer circuits.

Priming: Hybrid State Space Models From Pre-trained Transformers A mathematical framework for transformer circuits

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.614239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:6bb61ee4afecff14cc82d5d2963071a3c2ca912f6a4706e96579b0f005b6abbb

Observation 3cf975c1-395e-4bc3-9f16-7eedb95fba61 · outbound

This paper cites AREAL : A large-scale asynchronous reinforcement learning system for language reasoning.

Priming: Hybrid State Space Models From Pre-trained Transformers AREAL : A large-scale asynchronous reinforcement learning system for language reasoning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.603176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b9393fdfdc82e7d56bac2573c5a6faf074718c6ad743d1c9c10d1c9c3e69a390

Observation 426662d7-f5b7-4f1d-bd8b-15a064ad7429 · outbound

This paper cites arXiv preprint arXiv:2512.12167 (2025) 44.

Priming: Hybrid State Space Models From Pre-trained Transformers arXiv preprint arXiv:2512.12167 (2025) 44

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:24.341764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:33234a5a8a52ca0c6c31f4116741cff2ec66dfbf28715b36d73fe04adb415090

Observation 1cabd8dc-b39d-40ab-81f7-05a3ffbb71d6 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

Priming: Hybrid State Space Models From Pre-trained Transformers Zamba: A Compact 7B SSM Hybrid Model

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:24.308739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:3b24d154b60001dfe0d9d089ce4a59cacedee200c41a4cff36e78cc7d7671bbf

Observation ce0bfaac-6b8b-4da5-921e-5945e98fe755 · outbound

This paper cites RADLADS : Rapid attention distillation to linear attention decoders at scale.

Priming: Hybrid State Space Models From Pre-trained Transformers RADLADS : Rapid attention distillation to linear attention decoders at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.605008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:ce022f41823ebce84a7700433db24c07ae5f44bee2f4556093bb69d41bc84b4f

Observation f20a10db-f0f4-42aa-ac4f-4b9203b03277 · outbound

This paper cites The Llama 3 Herd of Models.

Priming: Hybrid State Space Models From Pre-trained Transformers The Llama 3 Herd of Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.292006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:3b8577d4378110c3c00545086626c97aa3ad29a8fcbe8e9c9df64e401d2f11ca

Observation 6ff6264f-8d7a-43a0-afb5-fd5dddbf66dd · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Priming: Hybrid State Space Models From Pre-trained Transformers Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.314225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:9cf09e810da39b6d5dc216ce4482c67fb2ecce3ba0ad8bd6f8fa9266713204e4

Observation 281c1068-8fdb-4e87-819f-1a4ac6d0763f · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

Priming: Hybrid State Space Models From Pre-trained Transformers Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.615947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f05bfce27e49890390fb0487260b12b1130a5d3b4e4393349461ebcec4c263fa

Observation e531da0d-4253-4a6d-bc90-599e5a9b4b53 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Priming: Hybrid State Space Models From Pre-trained Transformers Efficiently modeling long sequences with structured state spaces

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.632757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:366506df565fa9a6c2fb22cea6a9e0a60b5592cb454a76e03406650742c36f07

Observation 16c30fae-2d50-47a4-baec-22e9372604a0 · outbound

This paper cites Jet-nemotron: Efficient language model with post neural architecture search.

Priming: Hybrid State Space Models From Pre-trained Transformers Jet-nemotron: Efficient language model with post neural architecture search

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.634462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:835e66f552b8df3a8df4d2956e76ccd1da91364bc0d784311b640528ac44459f

Observation 3bf65d22-8989-4c53-97ea-5578daaa2961 · outbound

This paper cites A survey of model reduction by balanced truncation and some new results.

Priming: Hybrid State Space Models From Pre-trained Transformers A survey of model reduction by balanced truncation and some new results

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.599732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2cee1db21c1c9f4b444ee80dfcb25fc00f5c407d401da978a39746ce8ff5aeae

Observation 8c850831-affb-47d6-bb81-6e593771b881 · outbound

This paper cites Measuring massive multitask language understanding.

Priming: Hybrid State Space Models From Pre-trained Transformers Measuring massive multitask language understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.593886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:fd38f5c04474d6f3aaa5e3e0b889c8971af5079f7a079bf4463a9cd5ce9d126c

Observation f8cf7fb8-eb70-48aa-8284-7c31f7af2c0f · outbound

This paper cites RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling.

Priming: Hybrid State Space Models From Pre-trained Transformers RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.601514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:6033c61a497ac123e84a184cdc4158982a171ae70b8bd04f9ef53ec454f29ad1

Observation 1ae39118-3125-4b44-8c67-34f172dcf609 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Priming: Hybrid State Space Models From Pre-trained Transformers DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.650116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c6afe5405a4b3b8145a6b2f4bc7e7315b32363d8c5780182fa8fe602da691df0

Observation c30619b5-d517-4600-96a3-288bb4aaf694 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Priming: Hybrid State Space Models From Pre-trained Transformers Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.586669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c39b6d06bee3b874a119bf8e3aeeb508a670ca17fa959127e18f9c897f605de7

Observation e2ef0674-9eb1-41a4-aeda-305953802a16 · outbound

This paper cites Kakade, and eran malach.

Priming: Hybrid State Space Models From Pre-trained Transformers Kakade, and eran malach

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.606603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:8534ab096df6730af84d74a1fb227aa05657123b547e81a71aef8e7d0d1da94c

Observation 47c771e0-f2e4-4a40-bd43-0864906837c8 · outbound

This paper cites SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations.

Priming: Hybrid State Space Models From Pre-trained Transformers SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.581759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f9d4ff567d27509bb8651a495670c2942516801ffe7b576d9503500792a7c899

Observation 490f224e-a9ff-4fa4-88b5-de3383575342 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.583305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:e0e6430bb88e82f47424848812d2554e537666bacfabb044ce52de9e01916baf

Observation 2e58b95e-be58-4d01-9c25-cfc01b0401a0 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Priming: Hybrid State Space Models From Pre-trained Transformers Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.584966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b021ff0fb26849d1a3f524b129d218cf81450f0679f626352980dc539e0115fc

Observation 3a2a1bfb-6e45-4952-a779-28f2d39b268f · outbound

This paper cites Reformer: The efficient transformer.

Priming: Hybrid State Space Models From Pre-trained Transformers Reformer: The efficient transformer

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.588437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f268036b0fa8b7f02465bc2390b2ed0b403db4c760c8a457ab1bcc4032c5029a

Observation 47bb123b-44ac-4f48-944c-bf5a5a20f7b3 · outbound

This paper cites BABIL ong: Testing the limits of LLM s with long context reasoning-in-a-haystack.

Priming: Hybrid State Space Models From Pre-trained Transformers BABIL ong: Testing the limits of LLM s with long context reasoning-in-a-haystack

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.591885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:6796bd673bfeb6f94e2e7342618c0ee9c15cf397e18e10f1240dc6a48166f23f

Observation aea44e9c-66a0-4197-a0ae-7236307006a8 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Priming: Hybrid State Space Models From Pre-trained Transformers Gonzalez, Hao Zhang, and Ion Stoica

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.596001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:6d4df2e0c5a174273ad63f46f41b5ea699fb7d3f25c0fa2630810ce91808be8c

Observation f8e55d16-425b-49ef-8448-f5994b95eb4e · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

Priming: Hybrid State Space Models From Pre-trained Transformers Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.590105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2952c4a90694fdd1a1b7fc7c8ec531072e4855ca8f7f459e4c0dd9a4fb3ca87b

Observation 988de457-b21d-4390-a0c1-2fa629a902f9 · outbound

This paper cites Liger: Linearizing large language models to gated recurrent structures.

Priming: Hybrid State Space Models From Pre-trained Transformers Liger: Linearizing large language models to gated recurrent structures

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.579760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:ad6fec99ee90bf525418dd11e31e12d35c04e0d70abe1be97169ca6bf95c3033

Observation 28d1534a-5c58-4db3-ac42-84b8dc67bb81 · outbound

This paper cites Distilling to hybrid attention models via kl-guided layer selection.

Priming: Hybrid State Space Models From Pre-trained Transformers Distilling to hybrid attention models via kl-guided layer selection

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:24.433997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2185e3058d5b001f665d5446d33d052a16e178431dc5494bc509d04ff8d6e262

Observation a00fa614-b8a5-4108-a1d7-86958c0cce91 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Priming: Hybrid State Space Models From Pre-trained Transformers Jamba: A Hybrid Transformer-Mamba Language Model

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:11:27.316872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:45fe7e54be4e2649b3dfadc41bf841f22c63d3eebca1bd122ed1b4629cd52825

Observation 49ceeb95-dd09-4998-8c1a-9b81a21cbe86 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Priming: Hybrid State Space Models From Pre-trained Transformers Truthfulqa: Measuring how models mimic human falsehoods

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.569193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:cc3251ba7d69f921d3fb83d1ab6d16d63231fdcaf2ea3062683eba40f659ee26

Observation 5b1386f6-2a5e-4c76-bd15-7c542f625b06 · outbound

This paper cites On the stochastic realization problem.

Priming: Hybrid State Space Models From Pre-trained Transformers On the stochastic realization problem

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.570902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2116c2e9afe4827c21169410a2ef728dbcdefa0bc50b3f77046f4ca037b6bbae

Observation d8b7613b-0b3d-4d70-beda-a375c2958f6a · outbound

This paper cites Ringattention with blockwise transformers for near-infinite context.

Priming: Hybrid State Space Models From Pre-trained Transformers Ringattention with blockwise transformers for near-infinite context

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.572567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:bce2f3f371b200af24b5c6d3ba73120476cf9c250fb41bad858c2a395e299e0f

Observation 5188ea24-2971-405a-a5f6-d4d51d710546 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation.

Priming: Hybrid State Space Models From Pre-trained Transformers Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.574274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:95798967a06e090ca2223f4280d04460445388b92540bc90cf829c8dfc80a2e9

Observation 161b96d5-0ac4-425c-8728-ed406cbc5ae8 · outbound

This paper cites PICASO : Permutation-invariant context composition with state space models.

Priming: Hybrid State Space Models From Pre-trained Transformers PICASO : Permutation-invariant context composition with state space models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.563820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:97c15c60da4e0143e2b9ae229de162e3c6726976b7fc72cd3c328e1faccc04b7

Observation 565abb02-4b3a-4a55-8172-2a0bf9caab66 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.565525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:ce0e9d69f14e509577d8412de865301693a98f8ef8f3b09e16d5ebb9966327ea

Observation 06284579-59f1-4c7b-9b11-24059eedf42f · outbound

This paper cites Error propagation properties of recursive least-squares adaptation algorithms.

Priming: Hybrid State Space Models From Pre-trained Transformers Error propagation properties of recursive least-squares adaptation algorithms

Reference 55

Resolution
verified exact
doi, observed 2026-05-12T01:16:13.851121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:e19e9cf13870060f266451fc4b4feb005e51937d7f105440cf55ba82d3d98b27

Observation a4f5633b-649f-4f16-a69d-e599b577a8cd · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

Priming: Hybrid State Space Models From Pre-trained Transformers When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:33:08.656299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:0fba2c4b262d632b14939d713f7a500c79854fcf956e570606fd1b3d26ad1910

Observation 882de4f2-0ee7-421c-b007-f59a54d3144f · outbound

This paper cites AMC/AIME : MAA invitational competitions.

Priming: Hybrid State Space Models From Pre-trained Transformers AMC/AIME : MAA invitational competitions

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.576248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2bc3e66cc7919f1a353bf1a136334b387729199bd570e1aad7d8944af6130182

Observation df7df56e-506a-4215-8e78-6fb315c5543d · outbound

This paper cites Linearizing large language models.

Priming: Hybrid State Space Models From Pre-trained Transformers Linearizing large language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.560473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d92cafd052d1dce4082c07e5c3c1b7fbed1eadaefdf22804dca0ba53e3be09ce

Observation f562990e-7dfa-487a-bf7a-e7a8c67249af · outbound

This paper cites Landmark attention: Random-access infinite context length for transformers.

Priming: Hybrid State Space Models From Pre-trained Transformers Landmark attention: Random-access infinite context length for transformers

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.555534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:bcbcf734553d9cad482b890fcf432fe3325d2982894264e929546f96c094daa4

Observation e83cdf94-e5b8-4fbf-aaaf-6faf1ac9127b · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

Priming: Hybrid State Space Models From Pre-trained Transformers Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:17:00.308428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b02c1276e0e4c5c4a8ac973e1932540ecafc84299ffa4e8abb2e125fb64a64e4

Observation 3c0df41a-7846-4aab-9c6f-176c801b3f07 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.557174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:640e393f9d13ec5b59b5d51d27248d069fe24cba1ea1b09577852ab690ff5a0b

Observation 60347e00-26d8-4a50-ad55-ce99597ea7b9 · outbound

This paper cites Expansion span: Combining fading memory and retrieval in hybrid state space models.

Priming: Hybrid State Space Models From Pre-trained Transformers Expansion span: Combining fading memory and retrieval in hybrid state space models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.562085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:7030804a4f263c4719e8fc606377011a7a39605fb8cb2d0362533e5dfe83db8b

Observation c39bae94-aef0-409d-ac04-5f7881ed9c77 · outbound

This paper cites In-context Learning and Induction Heads.

Priming: Hybrid State Space Models From Pre-trained Transformers In-context Learning and Induction Heads

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.351717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:e46bbab428f7e9dea98ce9e3b91a2390c37e6d3757acbd16f174f312956425b7

Observation f5f4758e-6e05-4ffe-8dd6-90b914375b6f · outbound

This paper cites Resurrecting recurrent neural networks for long sequences.

Priming: Hybrid State Space Models From Pre-trained Transformers Resurrecting recurrent neural networks for long sequences

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.553782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:0b37d2ca3b1f5075a4dd82ac43461253b1817daa577fc58f2b11b6117430a038

Observation aee01451-f1bc-4d82-8a51-f6d9ca9c3f38 · outbound

This paper cites Marconi: Prefix caching for the era of hybrid LLM s.

Priming: Hybrid State Space Models From Pre-trained Transformers Marconi: Prefix caching for the era of hybrid LLM s

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.558885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d0d5ae7c3cdf347047638a379f9b269187697ac830883a22bf751447b0020f16

Observation dcf34717-e677-46d3-8a55-f08bc8ed5e90 · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

Priming: Hybrid State Space Models From Pre-trained Transformers Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.567478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:a5ff93a5b2ae2c37798ed80179921ac875d2f839c0bd07ff0c8064a05a0f3fad

Observation 272b97d0-e517-4bce-a3ea-da3b4cea7187 · outbound

This paper cites Time-Varying Systems and Computations.

Priming: Hybrid State Space Models From Pre-trained Transformers Time-Varying Systems and Computations

Reference 67

Resolution
verified exact
doi, observed 2026-05-12T01:16:13.867011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:3c45f2366bbc5d5294400c5a660f5730c39eca5e4a5420bcacc98da76731bbee

Observation 6436c4c4-57c5-4314-a932-a569b495373c · outbound

This paper cites Ya RN : Efficient context window extension of large language models.

Priming: Hybrid State Space Models From Pre-trained Transformers Ya RN : Efficient context window extension of large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.597985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:2b3c4ff7c092d037f8c20e406d41df2721974582e7ab18e011e59d4136a95b3c

Observation a8d63189-fa8e-4492-9239-41c59c345cb4 · outbound

This paper cites Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression.

Priming: Hybrid State Space Models From Pre-trained Transformers Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:04:17.792158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:e6f5b78f84c0f661ff81f6e0f8a251a207a25979dd3887d9544593eec5ce88b6

Observation 9fa358c9-7797-4344-8374-0f00905acf27 · outbound

This paper cites Generalizing verifiable instruction following.

Priming: Hybrid State Space Models From Pre-trained Transformers Generalizing verifiable instruction following

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.546893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b14b089077e2455a2f68f869da43405ceeb21fa59c8f4dd2d023156f484434c2

Observation a5ae0a2e-c954-49b4-9b55-9223e8962699 · outbound

This paper cites Ulysses sequence parallelism in the hugging face ecosystem.

Priming: Hybrid State Space Models From Pre-trained Transformers Ulysses sequence parallelism in the hugging face ecosystem

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.548622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:efd44474da5619914114668e12882ba93d9409d2a961ee6b018eb7a0b500eee2

Observation 10b9c153-478a-4296-b8cb-82b16c474648 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.543131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:4c121f7b4f9aba4130c542fa493d8d6e2423419978550413d5696dd2b1172e4d

Observation 6a65d74d-56b4-496f-bb0e-1a5633daa46f · outbound

This paper cites Samba: Simple hybrid state space models for efficient unlimited context language modeling.

Priming: Hybrid State Space Models From Pre-trained Transformers Samba: Simple hybrid state space models for efficient unlimited context language modeling

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.516151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f9d4674d3fc0a56706ea245843fa38c11e092123d6daec71129d7cdc80fa09ec

Observation 7ef43862-d7e3-4c5b-9a2e-a8a6ef7afcd8 · outbound

This paper cites Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi.

Priming: Hybrid State Space Models From Pre-trained Transformers Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:24.329790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:a765121609eb25e990a77bcb20f5eccb2493a648d4b25b05a89833ff0c8d1bda

Observation 13a2ee2d-c848-4f4c-859d-419c44e8051d · outbound

This paper cites Sandberg and A.

Priming: Hybrid State Space Models From Pre-trained Transformers Sandberg and A

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T01:16:13.875977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:b5c21c12b4b6a86c3a313b2fc516f53a7684d813934f93064c18c5ed14c2e171

Observation 9bd85cc4-e5dc-4a41-a84e-71edc2ef8922 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.544778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c2d591339bb8f03621047918dccf2b460ce29d87810e76698f0e280d32dca589

Observation dee51b14-2b7a-4c12-a90b-8989670958c9 · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.550215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:7002dded366cda3f9a374818fcc5c3ad6803b61de9bf4c1990f7d7486db41472

Observation 52ba9a51-c1c8-4f15-b6a8-98bc4b2fce2d · outbound

This paper cites Flashattention-3: Fast and accurate attention with asynchrony and low-precision.

Priming: Hybrid State Space Models From Pre-trained Transformers Flashattention-3: Fast and accurate attention with asynchrony and low-precision

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.521493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:54056efa649bb96729be7b360adf4d6a20b7b860c69ec7c076c6870efccbcd29

Observation 3b4d2990-9ea2-48e9-9dd1-197840eb6047 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Priming: Hybrid State Space Models From Pre-trained Transformers Hybridflow: A flexible and efficient rlhf framework

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:16:13.863162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:fc184005ab39a7a2cc44d97125f57176a03a45452d772195a4a4b4345c7a5422

Observation 326614f8-3d6d-40da-93d7-9b8f4ef33b72 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Priming: Hybrid State Space Models From Pre-trained Transformers Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.303495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:beb905976352e18ed44a49466f2492183945fa73c0feabca872595dfc23a18b7

Observation 8922d645-bb58-49af-94e9-c1af7a027842 · outbound

This paper cites Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation.

Priming: Hybrid State Space Models From Pre-trained Transformers Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.357221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:ebe34389adf99f1de505be143d9a72f7560605e7bab4f164a332c40f0340ac4c

Observation ec2e9b3b-4ffb-49a6-8046-950a486a530f · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

Priming: Hybrid State Space Models From Pre-trained Transformers Retentive Network: A Successor to Transformer for Large Language Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:21:24.346538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:a1b7f7840ef3c5993531a940221f24877351c86787b41b993ec431897e38b980

Observation b68dc4de-f625-4eae-8803-362998f0c8d6 · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

Priming: Hybrid State Space Models From Pre-trained Transformers Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.523437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f954e5e001fc76c8d89791d6b43779c99ce0e092fae50aa879d8dc0107bbff33

Observation 82074d65-a63b-463c-be2d-c2229ed33159 · outbound

This paper cites Scicode: A research coding benchmark curated by scientists.

Priming: Hybrid State Space Models From Pre-trained Transformers Scicode: A research coding benchmark curated by scientists

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.519793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:c99359877234832bc29a9624c6f3d07685a1b63754f5173b91704dce6e2b20d1

Observation cb3c5521-07db-497f-b58f-4e30e9442b2d · outbound

This paper cites Attention is all you need.

Priming: Hybrid State Space Models From Pre-trained Transformers Attention is all you need

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.518071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:fb9853bea1eefd772fc6312a570ddf0fb88125d0f1a62d4afcb68c9b9b76a500

Observation 66e2f336-3a1d-42fa-a17c-4159881331e2 · outbound

This paper cites Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries.

Priming: Hybrid State Space Models From Pre-trained Transformers Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:24.297719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d5f60fd56985b1faa79cc38e585e7b7c7d2d596867704a7954dcd2a168c0f73e

Observation bb8b7f7b-f4c7-4ec3-b938-630a75f25846 · outbound

This paper cites MesaNet: Sequence Modeling by Locally Optimal Test-Time Training.

Priming: Hybrid State Space Models From Pre-trained Transformers MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-06-04T02:07:39.171975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:4896a77dff3058e75c4429dd81a5327c0db1056a33e71a4089db77de5d282e61

Observation 98579f67-ec15-4123-ad7b-0588d2c509bd · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

Priming: Hybrid State Space Models From Pre-trained Transformers An Empirical Study of Mamba-based Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:31:04.038637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:30f36af13a1b83a12d775710b4b57901431d87fbb3cbab5048494a7002b18c6f

Observation 31c73377-1c84-489b-98cd-dae9d5eac527 · outbound

This paper cites The mamba in the llama: Distilling and accelerating hybrid models.

Priming: Hybrid State Space Models From Pre-trained Transformers The mamba in the llama: Distilling and accelerating hybrid models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.538336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:bdd908b7246599246fd393b92dd1869297b294fd31720d8b57c15430e1775b7d

Observation d67adf58-18a2-4995-9480-5d8e621832d2 · outbound

This paper cites M1: Towards scalable test-time compute with mamba reasoning models.

Priming: Hybrid State Space Models From Pre-trained Transformers M1: Towards scalable test-time compute with mamba reasoning models

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.540001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:75c71eeaca7ebb404b1e9daf1e5281932ce2dfad2f422b458fe962a315c8148b

Observation 5356abb7-9ac4-4a42-9bb1-bd0936c3871f · outbound

This paper cites an unresolved cited work.

Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-05-14T08:05:37.541613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:bd204000d5249a5869f282b5406e82a31786a5506be9a72fb1b3b881e07c673d

Observation 71612217-148e-4283-bb0c-1cfb27be1269 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Priming: Hybrid State Space Models From Pre-trained Transformers Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.551956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d70ffdfb4b78fd6528073269973f7d050ac82d0a430349e5e82fb5c64ab103f7

Observation 412068f2-b15f-4266-acb1-1420bb73e247 · outbound

This paper cites Effective long-context scaling of foundation models.

Priming: Hybrid State Space Models From Pre-trained Transformers Effective long-context scaling of foundation models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.533035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:f04edc1ed4228908ef9852d78271fbcb921e4bd1dc5eb49ea6d980b8f5fc54ad

Observation 8bc08899-e54e-4ec7-b1c7-61b869b9ce6e · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

Priming: Hybrid State Space Models From Pre-trained Transformers Gated linear attention transformers with hardware-efficient training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.534723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:d181e708e02f587dfbcfe832aaa8f82c4c5488396f55294c32eedda5247c1ace

Observation 8a41a46f-e508-4d72-8d45-09fdfb443138 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

Priming: Hybrid State Space Models From Pre-trained Transformers Parallelizing linear transformers with the delta rule over sequence length

Reference 95

Resolution
verified exact
doi, observed 2026-05-12T01:16:13.870615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:0cba64b72224ce0ec5aa144635488e0a9c60ddae9c38bdbdc6e0d448786fa56f

Observation 647af476-cda2-4fbc-abbe-77fa6825e88c · outbound

This paper cites Gated delta networks: Improving mamba2 with delta rule.

Priming: Hybrid State Space Models From Pre-trained Transformers Gated delta networks: Improving mamba2 with delta rule

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.577993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:068c006646051ead05749cc00eb8ee6ee45bc198481c16cd92038e73bf72dc75

Observation 0a75c586-7ea7-4804-82b0-6ca2995911c0 · outbound

This paper cites Ape: Faster and longer context-augmented generation via adaptive parallel encoding.

Priming: Hybrid State Space Models From Pre-trained Transformers Ape: Faster and longer context-augmented generation via adaptive parallel encoding

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.531221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:10808a8345ab00de2748c64aa7ef414f8c4017eb05e8cf1d6c31dbcbb85b89d1

Observation e6b20170-6593-4a3f-93b6-389b03d502ba · outbound

This paper cites Helmet: How to evaluate long-context language models effectively and thoroughly.

Priming: Hybrid State Space Models From Pre-trained Transformers Helmet: How to evaluate long-context language models effectively and thoroughly

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.529449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:70d48a60358b96fe46740418da10d5e5c8be80b22784541080835d7a60d02d13

Observation b03622f6-866b-47b2-848d-461d9e794b13 · outbound

This paper cites Native sparse attention: Hardware-aligned and natively trainable sparse attention.

Priming: Hybrid State Space Models From Pre-trained Transformers Native sparse attention: Hardware-aligned and natively trainable sparse attention

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T08:05:37.536451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:60efa44958d9d7200bfcb6582557c504907b13d966e1c1d63b5de31bae5f3cc5

Observation 4f4d1986-d94c-4de3-a448-3c909abf6af2 · outbound

This paper cites Stacked Residuals of Dynamic Layers for Time Series Anomaly Detection.

Priming: Hybrid State Space Models From Pre-trained Transformers Stacked Residuals of Dynamic Layers for Time Series Anomaly Detection

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:24.281364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-12T01:14:01.584159Z digest=sha256:43917ddb6ec5b95496b151fa22075e7fdce0149317dc2739d514d47cf390aa0f

Pith citing papers

Observation 6e95951f-c6f1-41fe-b0cd-289351c35abe · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Priming: Hybrid State Space Models From Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:39:20.915609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:39:20.915609Z digest=sha256:288809185bf15ad4ddf1a93e0fb9f95c333b2450452dcce4f2932bc1411b0845

Observation 7241869c-1acf-401b-a44d-95865c955a0c · inbound

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale cites this paper.

The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Priming: Hybrid State Space Models From Pre-trained Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T02:03:22.318771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:03:22.318771Z digest=sha256:ff6373a25d4511e7791850059b8e584b2f6c68704364623b6bd862bd0b6c9323

Observation 8d8a1feb-65df-46c1-ba5e-724df7bd875e · inbound

Memory for Large Language Models cites this paper.

Memory for Large Language Models Priming: Hybrid State Space Models From Pre-trained Transformers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:37:54.395898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:37:54.395898Z digest=sha256:63835fc0de40dd801ec4da14c63c32b22a7b877295c939729c54d6403179d60e