Pith. sign in

Paper Citation Record · LEDGER

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference

As of 10 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2502.02040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02040 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:40:56.623496Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4125d7d7-4910-4038-a307-f73154be051b · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.331392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.383105Z digest=sha256:0e39fc8a43fb8d2b499ec0ab9ab96747aa147ee6f5d21b5db692c85c029784e0

Observation d0f03468-926a-41b4-8a48-0838424117a1 · outbound

This paper cites Llm inference performance engineering: Best practices., 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Llm inference performance engineering: Best practices., 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.322083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.386863Z digest=sha256:b8265d96945813e2d9ef6f69cad1ce50172511336e1826cffe9323e36088c2ba

Observation 71a53602-126b-415c-b675-abe2ed4fdebc · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.390827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.390827Z digest=sha256:113d6bfdb5a7c41a09f5f5e32af3dc191f94df7813e5945414ce0d2da1f498ab

Observation 0fd5c2b9-a9c3-4750-aed2-84c18df135ad · outbound

This paper cites Colt5: Faster long-range transformers with conditional computation.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Colt5: Faster long-range transformers with conditional computation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.311780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.394133Z digest=sha256:16ae1257ef419e4466126f0b50b2545edf7fba9efcadcdf415aeeff996985046

Observation ba66cd9e-843e-44c9-bdbe-c1c10b38aae4 · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.302208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.397230Z digest=sha256:87a47e2761b5879dfde3d69faf4e584c7072d9fefc8f647843e614c323f716cf

Observation 6d74ad29-afe5-4099-95e7-84ca590ae6c9 · outbound

This paper cites Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.292785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.400215Z digest=sha256:2467cf3ed0bab141e15edfbab5d940b6df1f14ba2aca941cb3b5ef1d78ccfcf0

Observation 41fab238-730c-4342-91a8-007410e1ba14 · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:57.283289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.403333Z digest=sha256:57e2ae56b06b40767c5d639c8bed70595f499acef027301a6eac2ac4cebb687a

Observation b7497dd7-1fc9-440e-a26d-dc0aaff5af52 · outbound

This paper cites Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.273661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.406063Z digest=sha256:3118a5426a88d2bc870f421b6790da9173e293d2e873d7c7970daf1447a98565

Observation 0d7e73a3-b19b-4020-8b21-4c906ee9a86a · outbound

This paper cites Speculative Streaming: Fast LLM Inference without Auxiliary Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.409002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.409002Z digest=sha256:19fd203d088b8738829dc6b3f9a3b88039d1eb4c9174cd5651011ca23cdf24c8

Observation 0fd1e4e1-0fe1-44e3-ab53-3f16e10cdc7f · outbound

This paper cites Language models are few-shot learners.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.412320Z digest=sha256:29cf8744cd31a9f53adff88537b1721ca38a739976a8ff4913ec132b457ab523

Observation 4303cf81-200f-4f73-b870-0c06615b8231 · outbound

This paper cites Medusa: Simple framework for accelerating llm generation with multiple decoding heads.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Medusa: Simple framework for accelerating llm generation with multiple decoding heads

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.258058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.415675Z digest=sha256:6d2afc28ad4cdff88534a4a899cc4063be80463dd629d9dcbfeae71edc059e9d

Observation aac75849-a2df-4b96-a798-56abaf1bfb20 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating Large Language Model Decoding with Speculative Sampling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.418582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.418582Z digest=sha256:c7df65470e45dca9aca7497fe1c8bbb434ad3e49b2c885ecf2611d820c829d2a

Observation 130c8be1-1ea0-485f-89b6-1dc16780f1fd · outbound

This paper cites EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.421746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.421746Z digest=sha256:135eafe4ac4881fb7d10615f7dc37e525924312477a85cdebc9b04e65cbfa94d

Observation 8ac1876a-3b63-4d21-a25c-97cf08a7e166 · outbound

This paper cites DialogSum: A real-life scenario dialogue summarization dataset.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DialogSum: A real-life scenario dialogue summarization dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.248545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.425216Z digest=sha256:b618d1811415880c0ba1e208fb721ccdd2df6d5112e5d99b369ec4062a361a68

Observation 21915d66-c976-49e7-91c7-a3ebf8c8865f · outbound

This paper cites Koala instruction set documentation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala instruction set documentation, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.239004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.428185Z digest=sha256:07d1d0e4effd004fa04e9ffc81594a7d9fd2e1eaa5614a6420a300bbf4ab9771

Observation 227accd1-0895-45b8-bca0-66d1ef3531a3 · outbound

This paper cites Dai and C.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dai and C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.229623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.431458Z digest=sha256:646177339a18018372e4cb7cff25350a4adb24e457d86db0408b727265965965

Observation e26b19c1-73ca-4a02-901d-1d9b158496c8 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.434433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.434433Z digest=sha256:dbad74f72489def64552b0e8e08a20e5a30fb94de03bd83a1a095d9a3934e676

Observation e41a4539-c671-4783-bc6a-9b7406b46111 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.437585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.437585Z digest=sha256:09a841150e51cef8d6972902b4a612158b2cef24a7f06ae5383798a51b2247f6

Observation f6e487e9-a792-4216-8855-6e30d3dee94a · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.440954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.440954Z digest=sha256:374ea653bb85c7c2533983ea81b979e3b752409bc45200ab6177034aa63822cd

Observation 160c048d-2e2c-45f4-b846-0c76eae1e813 · outbound

This paper cites Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.219584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.443911Z digest=sha256:1165b77eee4ef3d8ba806e4d40421af3e50f929fcf9b489bbff9aa7490df0757

Observation 609d15ac-e549-4076-9b5a-8e8fab22c3dc · outbound

This paper cites Depth-adaptive transformer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Depth-adaptive transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.209682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.446767Z digest=sha256:91ed40f5f5d7e4b0913500b4ee6462da29deaa9da0952bc996ac25d6c8e812fc

Observation f0c36f0c-7c4e-4754-bc62-a18a23bbc08f · outbound

This paper cites Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:40:56.806396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.450038Z digest=sha256:ac73f3c7d1e4dbb4ae718b4c77c464ae87870102d596bb9b15ac266af604c164

Observation f9c0b924-8153-49ff-ba7b-e93db0e1bb26 · outbound

This paper cites Decoupled early time series classification using varied-length feature augmentation and gradient projection technique.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Decoupled early time series classification using varied-length feature augmentation and gradient projection technique

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.200580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.453137Z digest=sha256:8830a41308400f2f6d62c3269a5b7dae93d1712e63f3393c7aba389cc4b8cc55

Observation 7084d49a-8e22-478c-a066-6a00445acd05 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.191175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.456254Z digest=sha256:61e3762591e235baa4f1a07a60eabc5b7dad66b246fcd9470c5a05475f30ef79

Observation 4431a023-fc9e-4a15-b5d5-2dcf933839a3 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.182197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.459296Z digest=sha256:79431f792d8378534b6dc563f459b1e05bc99d64f9ebc62f77dafe08c63e2a84

Observation 1b7f3946-3ba4-45cb-9ffe-e1ee086430ec · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.462238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.462238Z digest=sha256:dc6ef85e76fc3660de3e789534519563ee5bf76d72280f6ee977818dda7830c7

Observation 6bb98c18-d580-4034-8917-807fbbb8ad18 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.465244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.465244Z digest=sha256:4c17500ddd678d2fb02512ee22ea9a878585b8e759225588432c3c5dcba39204

Observation 5d4d0187-f8ff-4d4b-b8aa-da7a922e888a · outbound

This paper cites Breaking the sequential dependency of llm inference using lookahead decoding, November 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Breaking the sequential dependency of llm inference using lookahead decoding, November 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.166407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.468436Z digest=sha256:caffb3eabb6ae3bd150fd4f3bedb82016be2b0b3611ae49f1f37691b7346288d

Observation 4aa6fb13-2c4e-41b8-aacc-e8cb3ca8c476 · outbound

This paper cites Garncarek and J.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Garncarek and J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.156345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.471577Z digest=sha256:e68f1c85d194ad8cc0b1e28fd655e559044d5382593323452f96e8b745251add

Observation ecfd5849-273a-42c5-98b4-890529b78c3a · outbound

This paper cites Koala: Dialogue-based fine-tuning improves factuality and safety of llms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala: Dialogue-based fine-tuning improves factuality and safety of llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.146814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.474505Z digest=sha256:a1b0f9b045bbc8947494a145f715aea167ec06e5f0e1123626b354b17c2a57d6

Observation 85ed8cbd-fb4c-4221-b2c5-350088de0f74 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference MiniLLM: On-Policy Distillation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.477410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.477410Z digest=sha256:0e9303607be41f6ea2a56ed6797e2ec23b0b8e2894fb350a3d50b0b3c110d82c

Observation bca9d024-aab7-43b6-9906-3b91c0ed9afd · outbound

This paper cites Identity mappings in deep residual networks.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Identity mappings in deep residual networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.136782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.480814Z digest=sha256:5bd99e1add9ca10e29bebca95f70c422ec57424ac2223380666e6cce2c390318

Observation 26d78d7a-9e40-4440-85c6-0dca8b7a41ed · outbound

This paper cites Dynabert: dynamic bert with adaptive width and depth.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dynabert: dynamic bert with adaptive width and depth

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.127174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.484066Z digest=sha256:4e591f093ba909283123b3758621db90791c7d8b035b8dec549576b47dad195b

Observation faac1eb5-92f3-46ec-9546-12b43bf53dec · outbound

This paper cites Adaptive mixtures of local experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Adaptive mixtures of local experts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.117441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.487126Z digest=sha256:af9914ccf38080f7617319b38c871a6e1425cc958891996ec15f1f3063e51c28

Observation 9f701b72-db33-40fa-87e4-9b0568cc5a4c · outbound

This paper cites Mixtral of Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixtral of Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.490622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.490622Z digest=sha256:007a8d4bbc98c7356877c2b5688160f199a3bdb6a6fcaff438a52de25faef987

Observation b1bca0b5-6a53-4789-89b7-97a58e7df821 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hierarchical mixtures of experts and the em algorithm

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.107682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.494312Z digest=sha256:f0211386029143b35c7f3f5605eeb4fddde4b61917199d128ab32745dd3e7e05

Observation 8fdaa0e2-c1c9-4eea-b6c6-0c14f2a38fa0 · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ten lessons from three generations shaped google’s tpuv4i

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.097080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.497877Z digest=sha256:939abe7e9c9a0b51238e5c93be216a4daa66e88e5cf662da23b1b52cd93485bc

Observation 5e1ca558-39b8-43f0-b235-0172eaaed0ea · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference In-datacenter performance analysis of a tensor processing unit

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.087012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.501446Z digest=sha256:451c464266d434e8c04a5c7a702d312d17913fd588fc3265937c17b79947d9c6

Observation 1e7dd89a-5e9b-4a86-9cba-04cb45604405 · outbound

This paper cites Gpus and the future of parallel computing.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gpus and the future of parallel computing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.076898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.504762Z digest=sha256:3828b562df5d5f311e91fafeafffe6598b3efd2a722c4d4b9cb8e3a64476c0de

Observation 1fe90847-784b-4f93-afd3-b61115c1f186 · outbound

This paper cites Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.067134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.508416Z digest=sha256:6079757b382dff87ae45d0f8fe8706dd7cfe68beec176511252aaf137e8076a5

Observation 178775d1-af33-4df4-9fab-7f32a86e552a · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.512100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.512100Z digest=sha256:77807d914260a5795634c284f7b5a2994c05a5010ba2aa4dcae67af4ce27bfdb

Observation c83a3b46-ca25-44a7-89b4-fdadeadc8250 · outbound

This paper cites GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.056698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.515875Z digest=sha256:b63fee2c6d07b32075d20ea9a88edadd11067dfa03fde272db81ad66e235182e

Observation bf0974c2-d4f8-4464-b951-dc8ea5df6241 · outbound

This paper cites Fast inference from transformers via speculative decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.519530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.519530Z digest=sha256:348674aa897340bf0fd1e78101d65a8bb2b3f706b96bd186b5afd28d56b7d623

Observation a989bc8e-992a-4d64-9e55-4178ae7d21e2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.523044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.523044Z digest=sha256:22d2df6d3d7d35e2885c5084a5a181c4c6720352794d84b61127faf9ac62611c

Observation a7e6870a-1e77-4fe2-9c32-194f4a17984c · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.526776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.526776Z digest=sha256:d4dc18ecd0f7ea5de1b96ca8b22800ac6414b5ae833aa260d8e7f8717960b070

Observation a878b6c1-c1d3-4107-8066-86a96b964406 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference OLMoE: Open Mixture-of-Experts Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.530396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.530396Z digest=sha256:0abb562935b323c65916a84191a9a95ec9db847b8283e4f30ea3e12833bc302a

Observation 7c104600-4e33-4238-8012-774f324cb833 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.041818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.533930Z digest=sha256:095fbce9f8c28549f5ec06cbe68f3b6e230e5f015df64d1f692669e496dc718e

Observation f9aa0612-9a4f-46d0-9529-450ac437ab24 · outbound

This paper cites Cuda c++ programming guide, 2021.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Cuda c++ programming guide, 2021

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.032278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.536413Z digest=sha256:cff455b96d91f8d2ef7374eb14af7980e6746b0bee862e9ce45869a98db9d7b1

Observation 6e7f2e82-0775-4321-b5de-c29ebf5ee8af · outbound

This paper cites GPT-4 Technical Report, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPT-4 Technical Report, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.538683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.538683Z digest=sha256:06f6085dccb04edec002dd13a19372748eafe50fb616ace27994b4e4405bbb37

Observation eaa65c5c-e415-45b0-a151-d99761e0f97d · outbound

This paper cites Language models are unsupervised multitask learners, 2019.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are unsupervised multitask learners, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.016826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.541014Z digest=sha256:6590496943a876d6552517d2e87eca755368da91bef13787c827128b229136dd

Observation 4b394c27-284a-410a-9876-ede0ab73677d · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.543541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.543541Z digest=sha256:63122b950099c7e0f2dae694a1721c2be4f9973e1798db9237c80a6cdf4f41ce

Observation 9ba1dab6-d238-47d8-b6e7-397f3f7d009d · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.546001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.546001Z digest=sha256:1807ca14fb3326779b559c2dea74c7f24763115e24a9e28e13bbd24dda8376da

Observation 99140b0c-0b9b-45c6-bccf-dd6f79f30650 · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tran, Yi Tay, and Donald Metzler

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.006694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.548559Z digest=sha256:a85be846513ec998f81e985c9a0eaa13c80708e1915427e0d0b1be52499dd3a9

Observation fc8612f8-4408-4a35-ba5e-20b685fb18db · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.551350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.551350Z digest=sha256:8e2a43b2324598b951b4999f9a8e729f30042255e2d12637e7dcea59325b0f6d

Observation 82d54c11-66cc-49ee-b739-1d0bc93339bc · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.995345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.555030Z digest=sha256:c9f0621e1c51bd3f79ff77379b89586944334a2d4486814244eda5f1c7b65ba2

Observation 88d5c12a-7214-4143-a264-5035e6bc7063 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.558179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.558179Z digest=sha256:a19a6a3d90c83b8cb6b086d29aa8fab0697617761b36fa70774e5bc5cef0155b

Observation f77c6337-9423-4108-8396-e2ff63b2e4f7 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating LLM Inference with Staged Speculative Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.561768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.561768Z digest=sha256:3be30cbdc6117f05ebb93537ed7712f57fa170267f63304ac38550950c9a1144

Observation 8082d3ff-2a22-46ec-b727-441fb821c775 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference A Simple and Effective Pruning Approach for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.565238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.565238Z digest=sha256:28c8e205a055da3a3baf53ab304861c82557edb76c066d1bc33d92b0160af24c

Observation 204e7fcc-a2aa-4a6b-a5f9-999fb832453c · outbound

This paper cites Manmatha.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Manmatha

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.985328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.569007Z digest=sha256:bdaf7dfd7496193058569286930734f859b45ef7b9b9946196fc729c87f8dc0b

Observation aab5ecd6-5ddd-4812-9a99-a9d0e01cbdea · outbound

This paper cites Tuan Pham, S.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tuan Pham, S

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.975515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.572295Z digest=sha256:fb128313ee815dde252fe63714f4275ea6343748325e233756c5cf790663b5bd

Observation 3340cb68-c9a3-433b-ac70-c5a9fc519aa3 · outbound

This paper cites Model cascading: Towards jointly improving efficiency and accuracy of nlp systems.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Model cascading: Towards jointly improving efficiency and accuracy of nlp systems

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.965723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.575573Z digest=sha256:4c0b8f5acb951e353d648ef0812fdef60c1e97a8b2e6963e838bd70ee7c66faa

Observation a84b56e0-c4d7-43ac-9b3e-6fa0effefe23 · outbound

This paper cites Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.955798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.578809Z digest=sha256:56c487cee0cad4ff086b527aeb451a9d802cbe90cf3fff2b4d975b4bc69c12a8

Observation eaccb3ce-68cd-4c3f-86a9-c4aa429352f3 · outbound

This paper cites Attention is all you need.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.581952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.581952Z digest=sha256:00d44729abb2f5574066cf6f7e22a816fe796beb292a0b68cd52b1a9750b14a5

Observation 941f8570-0355-4682-9b53-c23bb9ffce91 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.585245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.585245Z digest=sha256:a8503d254d949381101f02ff009663ecfb252a87d7ad8cd73ff802cd8d508487

Observation eb6434e6-a44b-44f0-8041-5c12645a8b5f · outbound

This paper cites Transformers: State-of-the-art natural language processing, 2020.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Transformers: State-of-the-art natural language processing, 2020

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.940331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.588725Z digest=sha256:f4159b4d90c6e83dbb985e31685fd9be0c283e7069d6568ad6389d296aed36c5

Observation f6407729-53b6-45dd-acbe-16cf625e36cf · outbound

This paper cites Speculative decoding: Lossless speedup of autoregressive translation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative decoding: Lossless speedup of autoregressive translation, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.930511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.591822Z digest=sha256:b0ffa1a9ed2c7ee467042305a27893b39c50d833032f44aee0908d4ddeb2ed7c

Observation 09a028c5-939f-4d78-ba5c-e2473c964c4d · outbound

This paper cites Deebert: Dynamic early exiting for accelerating bert inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Deebert: Dynamic early exiting for accelerating bert inference

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.920409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.595373Z digest=sha256:397d4c1819a5bb48e0fefe3044d41c6d7cc3a15d17c0521253b4d9b235af519f

Observation 772efeba-ed59-4ef5-a62b-88fe4ebbe676 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.598812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.598812Z digest=sha256:dd674185d72bbeb378a49f93b1b4e47b4965af89bd356b1d6f59e603b9351b34

Observation a3788be9-4e4b-4053-ad3c-d74088f4b257 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.909812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.602300Z digest=sha256:96d2ce04802f51be82c304fbf01f196dd302cd6dee58dd7bc42bd8110e698700

Observation 350e3e94-af68-4af6-9d1a-37eba339cc6f · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.605732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.605732Z digest=sha256:b7cfe9a84286d25cddbb3d2abe2df9e0e8ae5b7a0d6a0cde2d9dda7ad2e709f4

Observation e22e5582-7e5d-45ca-aeec-b77e78ccb127 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference P Xing, Hao Zhang, Joseph E

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.609457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.609457Z digest=sha256:1c1c84964ba8dadb320dc42b934b493b561f86e98a3a188b94ef1aab52808b83

Observation 334aef28-f2e3-4be0-b144-280af8aee97c · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.612810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.612810Z digest=sha256:656fb8c48e02a8243bd099eb3680f8f02914e6eb24f7526f5348682c349bcfa5

Observation da1245c6-485f-4294-88a2-171311cd026f · outbound

This paper cites Zhou and et al.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zhou and et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.893641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.616348Z digest=sha256:a0cb26fcc41308d75a609e31f91c4decceb6403e89b1219d09d0bd0acca4fc7a

Observation e687d854-34be-4b62-9d06-1176b7fdb6f1 · outbound

This paper cites Designing efficient sparse expert models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Designing efficient sparse expert models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.883584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.619675Z digest=sha256:fdd24049ae47c3e43e49f5447503788bc138962461f6f66bc8a897bd9ea4a5ca

Observation f72b9031-d6d4-4fbb-9136-bdaa7c543a1b · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.874916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T13:40:56.623496Z digest=sha256:2e41ea6d8517e7c997eedcd507dd53711472428facbc72ae68e6b00fd30488d0

Pith citing papers

No inbound Pith citation observations are available.