Pith. sign in

Paper Citation Record · LEDGER

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.22135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22135 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:31.702035Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:22:45.702572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T18:25:00.050125Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1d6a737-a5a7-4442-9199-b5d0eb899580 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding LongBench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:36.060030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:26.790813Z digest=sha256:351ba688a679305d7b0b3d6ec8f33272962290ad3865b69dcaf49852e6e7957d

Observation 3d0bfba0-6fbf-47d1-865a-ba680d1ea225 · outbound

This paper cites Puzzle: Distillation-Based NAS for Inference-Optimized LLMs.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:26.989808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:26.989808Z digest=sha256:568f5d7cdec83b42b615c4cd4b2cbf2d4e24114975764755f0151acd4b5bd7ad

Observation 7c9f6219-ef5a-41ae-815f-b24b6efb8679 · outbound

This paper cites On attention redundancy: A comprehensive study.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding On attention redundancy: A comprehensive study

Reference 3

Resolution
verified exact
doi, observed 2026-08-07T13:20:35.760682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:27.103525Z digest=sha256:b19fdee45ca4202781c19b3888968835cc1307fe3b5c4450b50aaaec84c36b88

Observation 37df9508-f6c9-4ea0-a92a-8dd431c00551 · outbound

This paper cites Xing, J Zico Kolter, and Albert Gu.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Xing, J Zico Kolter, and Albert Gu

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:35.505508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:27.248509Z digest=sha256:371eefb0a16631d1260f8009247bff1ac8b6028e5da3d86d9d929aab710a4547

Observation b38696a6-12dd-4325-b989-cdd99ed49a36 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.356000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.356000Z digest=sha256:a9e9f637211bea6ede81daec69995f39bd30e2b32987ed820cb40ab430966246

Observation ff0ac070-64ca-4102-9d9f-9fcba4d176da · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.450977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.450977Z digest=sha256:9224e0110effc87b8c29a7d6c1b850d681a76c15584b9b13603f492e2896d76b

Observation 363d877d-4e92-41a7-ac51-4bbce35aff6f · outbound

This paper cites GenQA: Generating Millions of Instructions from a Handful of Prompts.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding GenQA: Generating Millions of Instructions from a Handful of Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.547714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.547714Z digest=sha256:bf84beb52ba728b8bd8bf2e29e806ba2f2b265e18224b4670467d081ddcd56d9

Observation 59bef50d-e274-414d-ab60-f4a43793b7e6 · outbound

This paper cites Streamlining redundant layers to compress large language models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Streamlining redundant layers to compress large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.628183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.628183Z digest=sha256:31edcd35d9287eebafad60576234056b7cdc6f818f64841adf5c119a3ea5a89d

Observation 8753613a-08b4-4139-ba3d-9870820ea493 · outbound

This paper cites Stuffed mamba: State collapse and state capacity of rnn-based long-context modeling, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Stuffed mamba: State collapse and state capacity of rnn-based long-context modeling, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.719920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.719920Z digest=sha256:8880957f9174a1e54ea864bcb580b6c31c83cfa08f78df86d9ea41ec8df23821

Observation 8c5c7f89-39fa-4751-8e52-9cbae0213827 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.892094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.892094Z digest=sha256:512ce8f4cd0e05964f51a11c53aab45f421ae8e9c99a87b9dbc64c3b86dcdb0a

Observation 384e185f-0837-47e5-b562-9587a6e48c20 · outbound

This paper cites Analyzing redundancy in pretrained transformer models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Analyzing redundancy in pretrained transformer models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.023710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.023710Z digest=sha256:906fba811c3a401e70ee136b4447f87864ac7e5b57367ac55503f3a127edb904

Observation 583648ac-4682-4bdd-9aff-17ef39464327 · outbound

This paper cites Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:35.041855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:28.294265Z digest=sha256:d75d9f3c189461464f1a1ef3d20001e3a1f8ae9f31cf0009ee4c14fce35e5693

Observation cc73098e-1da3-4970-97b0-e7bf48a4461d · outbound

This paper cites Born again neural networks.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Born again neural networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.845967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:28.378302Z digest=sha256:67bf1424b52eebd7ed3925da957a5bbc302e09457d63f42bbebda739877f0bb1

Observation 71bc3d20-18fb-46f7-8624-67b907939de3 · outbound

This paper cites A framework for few-shot language model evaluation, September 2021.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding A framework for few-shot language model evaluation, September 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.465318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.465318Z digest=sha256:c3a8483984f62615a080d2977563e0db7df82a2f5f2b615fca719ef918565192

Observation 10680adf-502b-468a-802e-fa5f7ad8fe13 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Zamba: A Compact 7B SSM Hybrid Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.522002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.522002Z digest=sha256:a9cceb55a7e771422882bdcfc529938c2951be5de45fa2344cc5a87fb35e2bf1

Observation 6ae875eb-0f81-4bda-88bb-554d104e26fd · outbound

This paper cites RADLADS: Rapid attention distillation to linear attention decoders at scale, 2025.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding RADLADS: Rapid attention distillation to linear attention decoders at scale, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.654716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.654716Z digest=sha256:c7349296ef4c7ce9348dc5c65f80dfbd1bd99a8d496842f4a58ee9f8a32749c3

Observation 8c2edfd2-3690-4b7b-a3bf-b2f719c17cbf · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:34.597696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:28.773262Z digest=sha256:2371cf985914877de3dc4bab0f63ab439ce84cdee41fc813b18c34ea5bf32b7d

Observation 505045e6-a7dd-4d84-98fe-510186747f4a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.862191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.862191Z digest=sha256:39688ecccc59a80a6019c1ebd3871a8986deefe4ae65325ac65aed668d48b7c1

Observation 792ca933-5ff7-49ee-826d-caf3319ca0ea · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.974218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.974218Z digest=sha256:0c1779a73274a6da6426c6132d77b4efbdec2c764d0b454061071ff2601f539d

Observation 15bcccc1-c910-4cdc-baed-262a36cc83da · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Efficiently modeling long sequences with structured state spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.130322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.130322Z digest=sha256:55e73d35094812e10c4f1f874aad2a1f1d2f8b9d74937cdc0a6f5b06b8aef612

Observation 04ea52e9-2045-47f9-9246-077abb6e75e4 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.178468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.178468Z digest=sha256:61a68952c68778523abcd0815ac08cb76c2f41d5116baa58aa450734fcb26b44

Observation f7e90854-c69a-4f0a-a85f-233c2ebb1f1e · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding What Matters in Transformers? Not All Attention is Needed

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.224634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.224634Z digest=sha256:ba278332dc5a11a609e54c2c4e32c3bf5f34846f9b28c1d58429c83660ef61ce

Observation fc3a4e75-a77d-4ce1-9059-a8989c4d24dd · outbound

This paper cites Query-key normal- ization for transformers.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Query-key normal- ization for transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.313400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:29.315730Z digest=sha256:7af7ae6fe98c8837feb80e2f8684c3d3ffc9b3da93ddb39ed3a72322d50b6575

Observation abe84c62-093f-46ab-8521-57e4a50b91b5 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Distilling the Knowledge in a Neural Network

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.459790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.459790Z digest=sha256:8cdde2da4bc2dd1386ccdf4bd2035db671207d4302597615acd201cb68398e30

Observation e47c749d-e5bd-42bd-983c-2e4954533252 · outbound

This paper cites Kakade, and Eran Malach.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Kakade, and Eran Malach

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.069698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:29.576545Z digest=sha256:e00d50e60ba14cffde842ebc459625e95c631059d33b6bd6bda65a5bdf757af1

Observation 09b28262-294c-4137-b3e3-0ad36e27f9bb · outbound

This paper cites Fast inference from transformers via speculative decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Fast inference from transformers via speculative decoding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.771256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:29.652714Z digest=sha256:d6fcb5feabd73c1dfc143e2f5b4d62d3f63625f6665fa0a513e674633dffde91

Observation 75c86232-ff5e-4dd8-a5a0-b90bd121581f · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Jamba: A Hybrid Transformer-Mamba Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.796292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.796292Z digest=sha256:e5e052b805b9bfad07acb2d79f2b7d2d82a66aa830b6193f568e998759c1eac5

Observation dda98709-d7ab-4e42-859f-828c0bf10a46 · outbound

This paper cites ZeroEval: A unified framework for evaluating language models, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding ZeroEval: A unified framework for evaluating language models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.495066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:29.900421Z digest=sha256:c432632ad2dd00fbaa0a5e3ee2d50a7f4e4ad1d0c1303c4c8bd245e4d4548544

Observation f53e4d36-c4fd-4546-b2f3-19cc08572a2b · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Longhorn: State Space Models are Amortized Online Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.988143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.988143Z digest=sha256:ccbd09771d20faa35f65e22642a68e44fd1aa0b50d5ebc08d9b0d58d7a8803fb

Observation 9ad9e62b-be82-4353-ae17-d2abbba95ad4 · outbound

This paper cites ShortGPT: Layers in large language models are more redundant than you expect, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding ShortGPT: Layers in large language models are more redundant than you expect, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.250157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.103410Z digest=sha256:f697f88f912b058668230d7acb5ae378b07a4c1820cf315827e19d839364bed3

Observation 889c22b1-ed36-471e-a1aa-f6dedd4dd1b1 · outbound

This paper cites Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32, 2019.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.150611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.170022Z digest=sha256:553fdd0d6bfd3a7b32c43a3a49121454b3b9417c3cc6e6e59df00b5ef500a70c

Observation 7f12d9e6-c704-420a-933c-0f6509bec4c7 · outbound

This paper cites Compact Language Models via Pruning and Knowledge Distillation.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Compact Language Models via Pruning and Knowledge Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.255508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.255508Z digest=sha256:39d90554ac92c625af297087ee2969b2d2256e21d153a228f9fcedb9d805f72c

Observation 9ee686fd-fb30-4a8d-ae89-6a29e4acba5a · outbound

This paper cites Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.044723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.313713Z digest=sha256:9564f5c7f4849be4d86717fc41f19ed237c0cdf765abc40bc3152b3643a0eee8

Observation d6b35deb-f0e0-445e-9e10-a7ed99dbd88e · outbound

This paper cites Infinity instruct.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Infinity instruct

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.936382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.359409Z digest=sha256:d40aee4e3502ef1f9ae9efee0c3cd5a9b65c5a1507cf9269f5379a6226069294

Observation 72ced6b8-b35b-453a-b537-264ad8ee29d0 · outbound

This paper cites RWKV-7 "Goose" with Expressive Dynamic State Evolution.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding RWKV-7 "Goose" with Expressive Dynamic State Evolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.438474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.438474Z digest=sha256:ec51faf8e7fe9e1ff71ac5975a8e5f768c2b88314ab93838b3df2bd3c1f4ee5c

Observation 019756d4-33e6-40b8-8be0-e74ef74e8d78 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Compressive Transformers for Long-Range Sequence Modelling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.491759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.491759Z digest=sha256:7eafde318ba936eb048630260537eb192c0e8109d1fe79b90473996ea1cb7c64

Observation f1eaacc1-e7b1-4749-876b-ef971285a3f2 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.608844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.608844Z digest=sha256:d27a09d8c2d2eea5f70456584b394123c573d4ab4661c6690e3dbf66e9fc6b94

Observation cea3d994-b524-4a3b-b4c1-63a1a747d507 · outbound

This paper cites TAID: Temporally adaptive interpo- lated distillation for efficient knowledge transfer in language models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding TAID: Temporally adaptive interpo- lated distillation for efficient knowledge transfer in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.852697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.666488Z digest=sha256:7b6e6027af55cdc87e52130d57a51176063e26de494dfb76d93521c4a7d78d8d

Observation ec1d226a-ee4e-4caa-8427-5d7dc4317c79 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.755203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.755203Z digest=sha256:7971918d176a57e3bc39c4ba61c6b1dbd5e5dfc3affcf94d22b4b089a6aaa08c

Observation 610db4f9-7003-4e02-a179-f5376252a4f0 · outbound

This paper cites Transformer Layers as Painters.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Transformer Layers as Painters

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.828566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.828566Z digest=sha256:868c0965274d86522042600afa5f44ac3f8e0aaf5358de923492df9729628c87

Observation ab900abb-5524-4b8b-8cf8-d262fa74d6a1 · outbound

This paper cites OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants, 2023.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.764491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:30.967987Z digest=sha256:05c771d548a56d07678e57a9deb20490634d300efeb3b1d2db5d023d38d9aea4

Observation b904818a-7db9-4b9f-9bc3-172616b79537 · outbound

This paper cites Attention is all you need.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Attention is all you need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.059850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.059850Z digest=sha256:c6862b2aace54c7feec6fd87ddafb1c521bb1d52960d58ac9bc1aad2c23306ed

Observation 663bb92a-f0a8-4a5e-95a5-02ed4cb3d922 · outbound

This paper cites Rush, and Tri Dao.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Rush, and Tri Dao

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.646348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:31.148458Z digest=sha256:cc2b764f3ed73eeee5b46af18f50d6cd5e6520f3caa82810fcbb5ebdefaae23e

Observation 35d6f519-5f1e-4140-8bb9-330f78d21323 · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:32.538808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:31.282968Z digest=sha256:21e21da54b530bc890700a870f21c86c4f29573ff1a899f2fa2f4d4608b91298

Observation 03735d46-1613-48b5-970a-3f4fed6f6db8 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Parallelizing linear transformers with the delta rule over sequence length

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.403849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:31.400726Z digest=sha256:e20479ba63e6a4e2813d540a91fe2e1510860f386d78205fd050bbaacc6caadc

Observation 1aa69c5a-6049-4b7c-bf0f-4ca029f51e12 · outbound

This paper cites Gated delta networks: Improving Mamba2 with delta rule.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Gated delta networks: Improving Mamba2 with delta rule

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.287914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:31.483657Z digest=sha256:3abc9b9dd1590788b7245725514f5921ceedca79402d74ebce02cc5711f00dc4

Observation f5f05cd6-5fe4-4479-b4c4-e5edecafde18 · outbound

This paper cites KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.147055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:31.572565Z digest=sha256:7595779a2b81902ac035a042dfcbde9fcad834cbbe09461f6f26d17d33255dfd

Observation 893308ac-ba7d-42ad-97ec-976addc14d2c · outbound

This paper cites Root mean square layer normalization.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Root mean square layer normalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.661089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.661089Z digest=sha256:377c15821bd0750b20f56db889e4bb6077d7fe2359a59a4c5919c05d75ff4f0c

Observation 2bdbaa1d-2463-486a-abd3-f3156652c70c · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.702035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.702035Z digest=sha256:3b891172860a45aea4f4e1bc2efcc638291c95445a7c939610299268e6638007

Observation 06fca4ef-0e7b-4b0e-b5cd-da60bcd63155 · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 398

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:35.277833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:20:28.145383Z digest=sha256:d2de9ad5ef749ee68de41dd14f6ca7cdc324285df31e3a2c7d64855c0927e3e9

Observation b9a285ad-0920-4b57-a372-21208309bb95 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.172.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding doi: 10.18653/v1/2024.acl-long.172

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:26.876727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:26.876727Z digest=sha256:5b3c5cc8b3b982f6764832edb43e67d152c643b1e840a5268059504cc2d50c43

Pith citing papers

Observation 1979501e-7fdf-4a9e-a710-451b31732a64 · inbound

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling cites this paper.

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:10.802659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:39:37.485602Z digest=sha256:78e3f1272f67fabbc4c0e0154c4da10137cb81b7d882dfa66b4a9190ab3ce44a

Observation 697489d4-87b9-4603-9952-c15806008e35 · inbound

Component-Aware Self-Speculative Decoding in Hybrid Language Models cites this paper.

Component-Aware Self-Speculative Decoding in Hybrid Language Models RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:40.749250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T18:40:17.017805Z digest=sha256:13198cc5306dc5d107b5a47a3f1062f7aff2a2679bb75a0a1c9458ad94db100c

Observation bd2fc1fd-38a6-409a-8940-04dabbe38e7e · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.184871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:00:35.496822Z digest=sha256:64a4df151fb6df6dbaf2559633b3fe0419ec72937662276197371e805c4b271d

Observation 6c0ae147-ab2a-483c-9629-e5dca0897884 · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:25:00.051504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:22:45.702572Z digest=sha256:6d3ee6d680a322c55e462722468a3a4eb02be1ccba496fd1e39e88cf0e2f4b17