Pith. sign in

Paper Citation Record · LEDGER

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference

As of 22 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 0 inbound Pith citation observations for arXiv:2502.02040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02040 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:40:56.623496Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact1
  • verified fuzzy41
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4125d7d7-4910-4038-a307-f73154be051b · outbound

This paper cites Phi-3 technical report: A highly capable language model locally on your phone, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Phi-3 technical report: A highly capable language model locally on your phone, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.331392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.383105Z digest=sha256:77f71842890ec2f8a239d7e8d889c81c23caf8354d39c7c76576f6eb250420a3

Observation d0f03468-926a-41b4-8a48-0838424117a1 · outbound

This paper cites Llm inference performance engineering: Best practices., 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Llm inference performance engineering: Best practices., 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.322083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.386863Z digest=sha256:aac5dd3ee0ae42d2cd01b591591ec479954f9cc2f42a64cf89771cd72b14438b

Observation 71a53602-126b-415c-b675-abe2ed4fdebc · outbound

This paper cites On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.390827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.390827Z digest=sha256:4b939a313d2735300832d6489004ca59863b67a0bcb08e5ff39fd1490aeeaf39

Observation 0fd5c2b9-a9c3-4750-aed2-84c18df135ad · outbound

This paper cites Colt5: Faster long-range transformers with conditional computation.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Colt5: Faster long-range transformers with conditional computation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.311780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.394133Z digest=sha256:5db9cde0e1543e9d0c3544418f089c1a1e3d2e0a23eb5500d7363a4c1b8223c7

Observation ba66cd9e-843e-44c9-bdbe-c1c10b38aae4 · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hydra: Sequentially-dependent draft heads for medusa de- coding, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.302208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.397230Z digest=sha256:710b459f6072bf0c4ada990717080e450f83632bc389effb34a2ca51aa7a1003

Observation 6d74ad29-afe5-4099-95e7-84ca590ae6c9 · outbound

This paper cites Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Massively multilingual sentence embeddings for zero- shot cross-lingual transfer and beyond

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.292785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.400215Z digest=sha256:a1828973581b469a4484755840c0b38573e53cac091a6312b085a47779923eb2

Observation 41fab238-730c-4342-91a8-007410e1ba14 · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:57.283289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.403333Z digest=sha256:ae66b3f8b496f4a9533736325799f5e880ea889e7a64d6f49dae97c8ba8e28fd

Observation b7497dd7-1fc9-440e-a26d-dc0aaff5af52 · outbound

This paper cites Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mt-bench-101: A fine-grained benchmark for evaluating large language models in multi-turn dialogues

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.273661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.406063Z digest=sha256:bcb98b45b06c7dfe6a446dd7d3543a42a716a934dadbc11239f9430e45e61bea

Observation 0d7e73a3-b19b-4020-8b21-4c906ee9a86a · outbound

This paper cites Speculative Streaming: Fast LLM Inference without Auxiliary Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative Streaming: Fast LLM Inference without Auxiliary Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.409002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.409002Z digest=sha256:fa0a799b60373d0b3619a5a1e4aa4d3ee9cbc1f97e89932be8ef2bc30043d8ae

Observation 0fd1e4e1-0fe1-44e3-ab53-3f16e10cdc7f · outbound

This paper cites Language models are few-shot learners.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.412320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.412320Z digest=sha256:bb7fa2ac2ec583ed55b23ce88fbb2682bef97e1a90929bfe203c01e65b7b31d2

Observation 4303cf81-200f-4f73-b870-0c06615b8231 · outbound

This paper cites Medusa: Simple framework for accelerating llm generation with multiple decoding heads.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Medusa: Simple framework for accelerating llm generation with multiple decoding heads

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.258058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.415675Z digest=sha256:339d99bc06916c1a46cfa56e748f60ab0070544f7bde6db92d1cf26a3030ac5d

Observation aac75849-a2df-4b96-a798-56abaf1bfb20 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating Large Language Model Decoding with Speculative Sampling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.418582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.418582Z digest=sha256:75f701bab24de2f9231aca4535e934fcd1f59625fd854592e904e70e94f23f1f

Observation 130c8be1-1ea0-485f-89b6-1dc16780f1fd · outbound

This paper cites EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.421746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.421746Z digest=sha256:a80107a3317039e67ccf02b799e0c992072bccfecbdd1b990283d8c0ec9086a9

Observation 8ac1876a-3b63-4d21-a25c-97cf08a7e166 · outbound

This paper cites DialogSum: A real-life scenario dialogue summarization dataset.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DialogSum: A real-life scenario dialogue summarization dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.248545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.425216Z digest=sha256:d4828775da5425f7dddbb2015f9000e22958724a66d3c61ac996f56d5b5c13ec

Observation 21915d66-c976-49e7-91c7-a3ebf8c8865f · outbound

This paper cites Koala instruction set documentation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala instruction set documentation, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.239004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.428185Z digest=sha256:e8516b5aca429be934db10d6b6aa04bea5d14a1c3637d7d0292c89c92519f7a0

Observation 227accd1-0895-45b8-bca0-66d1ef3531a3 · outbound

This paper cites Dai and C.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dai and C

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.229623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.431458Z digest=sha256:d2ab35c7e1bb55acbd263872fd4a7391df17fa87a6777ccf5185d75beee3219b

Observation e26b19c1-73ca-4a02-901d-1d9b158496c8 · outbound

This paper cites SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SkipDecode: Autoregressive Skip Decoding with Batching and Caching for Efficient LLM Inference

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.434433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.434433Z digest=sha256:10a462bbd195b14aef8063277279b27807bba59facd7c3eacca270b0e551b4c8

Observation e41a4539-c671-4783-bc6a-9b7406b46111 · outbound

This paper cites SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.437585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.437585Z digest=sha256:aa315866f9ee9df1111f09daa9709e03eb1baaffc6ac5b0ac102695d027a201d

Observation f6e487e9-a792-4216-8855-6e30d3dee94a · outbound

This paper cites GLaM: Efficient Scaling of Language Models with Mixture-of-Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GLaM: Efficient Scaling of Language Models with Mixture-of-Experts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.440954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.440954Z digest=sha256:18238566b4889911e78ffedd6a51a7badb9f474a47a5a7657714d49fb437485c

Observation 160c048d-2e2c-45f4-b846-0c76eae1e813 · outbound

This paper cites Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.219584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.443911Z digest=sha256:cf657958cc6ba46c63fe8a376311216f282d1bb8374c25917c0a582566c91242

Observation 609d15ac-e549-4076-9b5a-8e8fab22c3dc · outbound

This paper cites Depth-adaptive transformer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Depth-adaptive transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.209682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.446767Z digest=sha256:fce5c9cb56478fc925416130905db20fdd2a230a51e9f7cddcdfec234fa5b98e

Observation f0c36f0c-7c4e-4754-bc62-a18a23bbc08f · outbound

This paper cites Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Predictive Exit: Prediction of Fine-Grained Early Exits for Computation- and Energy-Efficient Inference

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:40:56.806396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.450038Z digest=sha256:6fa9842197c0c3cd7673875056f9b40aa0d1c76b7f27a3717bef478134433c25

Observation f9c0b924-8153-49ff-ba7b-e93db0e1bb26 · outbound

This paper cites Decoupled early time series classification using varied-length feature augmentation and gradient projection technique.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Decoupled early time series classification using varied-length feature augmentation and gradient projection technique

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.200580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.453137Z digest=sha256:f2dc3a5a2f9a588416b3b925323d0cd01742c6a2bd5b21ec0ba0537b6f8070a2

Observation 7084d49a-8e22-478c-a066-6a00445acd05 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.191175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.456254Z digest=sha256:0746d7254b09181bad3eb87539d568b05e8b0c8dfa313c280f480b3b950ec059

Observation 4431a023-fc9e-4a15-b5d5-2dcf933839a3 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.182197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.459296Z digest=sha256:126edc56036f1ab1f9355c46dd200199347c34de481903f82bfd4eae863fd618

Observation 1b7f3946-3ba4-45cb-9ffe-e1ee086430ec · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.462238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.462238Z digest=sha256:5b5e1be1e2735bea05d121001c271f792262db692beca52b280da0ddcfa81339

Observation 6bb98c18-d580-4034-8917-807fbbb8ad18 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.465244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.465244Z digest=sha256:23cd31e0fa88e7b060b16631be479751279a52c000cadeff8928fe8d26e28dc5

Observation 5d4d0187-f8ff-4d4b-b8aa-da7a922e888a · outbound

This paper cites Breaking the sequential dependency of llm inference using lookahead decoding, November 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Breaking the sequential dependency of llm inference using lookahead decoding, November 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.166407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.468436Z digest=sha256:71bc5fb02cca0cb3c6c215602f35b0e01956fd69efee3be861d42203184d558a

Observation 4aa6fb13-2c4e-41b8-aacc-e8cb3ca8c476 · outbound

This paper cites Garncarek and J.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Garncarek and J

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.156345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.471577Z digest=sha256:98068fd2279a1b59424594f75b2225b841b4735165f16fbf882d335c85beb56f

Observation ecfd5849-273a-42c5-98b4-890529b78c3a · outbound

This paper cites Koala: Dialogue-based fine-tuning improves factuality and safety of llms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Koala: Dialogue-based fine-tuning improves factuality and safety of llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.146814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.474505Z digest=sha256:3378da6f1b0ecf3df8c9b8998a4f5591eea502b2fc0299c29ea58dfc2121b1f5

Observation 85ed8cbd-fb4c-4221-b2c5-350088de0f74 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference MiniLLM: On-Policy Distillation of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.477410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.477410Z digest=sha256:4aebd37f05ba6e0ecc1a04a44f20bf8424216a5ea3604db2b5053a039dd1c533

Observation bca9d024-aab7-43b6-9906-3b91c0ed9afd · outbound

This paper cites Identity mappings in deep residual networks.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Identity mappings in deep residual networks

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.136782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.480814Z digest=sha256:8984ff8fd772f7816178a5e27b94ebe55928847ae11112c4ec667bac45527c69

Observation 26d78d7a-9e40-4440-85c6-0dca8b7a41ed · outbound

This paper cites Dynabert: dynamic bert with adaptive width and depth.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Dynabert: dynamic bert with adaptive width and depth

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.127174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.484066Z digest=sha256:6bc4bd94c4100421797e500316101b680f4bf9db511fb162d3079d4600a81be2

Observation faac1eb5-92f3-46ec-9546-12b43bf53dec · outbound

This paper cites Adaptive mixtures of local experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Adaptive mixtures of local experts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.117441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.487126Z digest=sha256:71b31a58dc1feb331767fb0e95bed35c6ca0ae12a0cd64bbf27b43dc24bbb6b2

Observation 9f701b72-db33-40fa-87e4-9b0568cc5a4c · outbound

This paper cites Mixtral of Experts.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixtral of Experts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.490622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.490622Z digest=sha256:e3369d67b9616f633a83b53c06b62fcca75b7535db6930d95e0663704e754afd

Observation b1bca0b5-6a53-4789-89b7-97a58e7df821 · outbound

This paper cites Hierarchical mixtures of experts and the em algorithm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Hierarchical mixtures of experts and the em algorithm

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.107682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.494312Z digest=sha256:0f0ac353944bb1180d357e68bd7f691d81ac6158ee3f754638e42b65829d8e8f

Observation 8fdaa0e2-c1c9-4eea-b6c6-0c14f2a38fa0 · outbound

This paper cites Ten lessons from three generations shaped google’s tpuv4i.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ten lessons from three generations shaped google’s tpuv4i

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.097080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.497877Z digest=sha256:05209c3f526e0f50566fe314e0e576799a3e8b65b1dcc5f6d6a4682a70acfc0b

Observation 5e1ca558-39b8-43f0-b235-0172eaaed0ea · outbound

This paper cites In-datacenter performance analysis of a tensor processing unit.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference In-datacenter performance analysis of a tensor processing unit

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.087012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.501446Z digest=sha256:a24aed20c5882f03682ae50db477b8bedf6c451c4a87ac6c7394b1f4c787f1be

Observation 1e7dd89a-5e9b-4a86-9cba-04cb45604405 · outbound

This paper cites Gpus and the future of parallel computing.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gpus and the future of parallel computing

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.076898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.504762Z digest=sha256:6126895e615f972f3619098518a6cbb1d4ec880952bef9fde03ee5a986f47bd1

Observation 1fe90847-784b-4f93-afd3-b61115c1f186 · outbound

This paper cites Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Ai and machine learning acceleration in mobile devices: A survey of architectures, hardware, and algorithms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.067134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.508416Z digest=sha256:253d549c8ccdc85c87739b9479021c2722ce2aea7c2e5735a7a27d67153e0a75

Observation 178775d1-af33-4df4-9fab-7f32a86e552a · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.512100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.512100Z digest=sha256:cfe67b9f69713ebb3055bda34e5e2649fc78c4013c1f47d50cdb4256442c78a6

Observation c83a3b46-ca25-44a7-89b4-fdadeadc8250 · outbound

This paper cites GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GShard: Scaling Giant Models with Condi- tional Computation and Automatic Sharding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.056698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.515875Z digest=sha256:16968c9c2d9c64d8b825df22b5fa657bb12ef9287ae557249a5a3df0d647d5f4

Observation bf0974c2-d4f8-4464-b951-dc8ea5df6241 · outbound

This paper cites Fast inference from transformers via speculative decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Fast inference from transformers via speculative decoding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.519530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.519530Z digest=sha256:8b10587f054669e2ecaa5906403838c6158dce1ab351e233658c7e721629892e

Observation a989bc8e-992a-4d64-9e55-4178ae7d21e2 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.523044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.523044Z digest=sha256:e47dca8a42df76aeee3981a8a9ffb8ada35fce676085ae03095e762887a60816

Observation a7e6870a-1e77-4fe2-9c32-194f4a17984c · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.526776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.526776Z digest=sha256:e2025e00c29406e23a4665ba340b038629b841922c6e5a16d54be8b791c714da

Observation a878b6c1-c1d3-4107-8066-86a96b964406 · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference OLMoE: Open Mixture-of-Experts Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.530396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.530396Z digest=sha256:4d6d25dbd8d3b5ee1b6cc96dc9ff3fe4e01b55ebed0b9e7d3c70c47369f10933

Observation 7c104600-4e33-4238-8012-774f324cb833 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.041818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.533930Z digest=sha256:6d5cf2fa50c14f8eed3c2c430360cf35dab664761f5a238d4ee9c85e988865a8

Observation f9aa0612-9a4f-46d0-9529-450ac437ab24 · outbound

This paper cites Cuda c++ programming guide, 2021.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Cuda c++ programming guide, 2021

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.032278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.536413Z digest=sha256:78d5937be222c219ed3789daf6711ba0d0d3ab238ba7e2dc31b408cadcc071e3

Observation 6e7f2e82-0775-4321-b5de-c29ebf5ee8af · outbound

This paper cites GPT-4 Technical Report, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference GPT-4 Technical Report, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.538683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.538683Z digest=sha256:80214f1fa0f1a633ddac85c2e0c9f9f2ec19b5d3d2cb940ed5e2151262f76cfd

Observation eaa65c5c-e415-45b0-a151-d99761e0f97d · outbound

This paper cites Language models are unsupervised multitask learners, 2019.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Language models are unsupervised multitask learners, 2019

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.016826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.541014Z digest=sha256:02b33a88378f46d78ed03a403d21445105e85eb0c407f1dda4d70971d5016092

Observation 4b394c27-284a-410a-9876-ede0ab73677d · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.543541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.543541Z digest=sha256:e3320982da7a1146a3f291e057feb712d68fb33a08dd957f3f8de7e97ecb6562

Observation 9ba1dab6-d238-47d8-b6e7-397f3f7d009d · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.546001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.546001Z digest=sha256:863048c498d77b808a5d6ff7c220275a92b956385f80b340b19196c8e14f9aed

Observation 99140b0c-0b9b-45c6-bccf-dd6f79f30650 · outbound

This paper cites Tran, Yi Tay, and Donald Metzler.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tran, Yi Tay, and Donald Metzler

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:57.006694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.548559Z digest=sha256:22969e42fcc291c24d4b80084f3d6a565b387835b643f4c3b1c11fe5d2b9191e

Observation fc8612f8-4408-4a35-ba5e-20b685fb18db · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.551350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.551350Z digest=sha256:e9bea149ef9f865bd9fa91f52ee36a9fd3abcc91d5c534b41ee1d251105e5079

Observation 82d54c11-66cc-49ee-b739-1d0bc93339bc · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.995345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.555030Z digest=sha256:5f54dbf451110f63c7ea6599986b478d006aa96777eec3d15148b75bfb16871b

Observation 88d5c12a-7214-4143-a264-5035e6bc7063 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.558179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.558179Z digest=sha256:7dfd6871cc39544cb770d5e9efb0a9f59ed0c488d5c99db9a2460899277cb088

Observation f77c6337-9423-4108-8396-e2ff63b2e4f7 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Accelerating LLM Inference with Staged Speculative Decoding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.561768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.561768Z digest=sha256:47c99d44e5f879c745ca665c8391d1adf01d89e6aa80eef166cbb86420a975b2

Observation 8082d3ff-2a22-46ec-b727-441fb821c775 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference A Simple and Effective Pruning Approach for Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.565238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.565238Z digest=sha256:710ea032fc84dee5e76aba0bda6c9634b6f89771b03d2e108b52e938269ed6d3

Observation 204e7fcc-a2aa-4a6b-a5f9-999fb832453c · outbound

This paper cites Manmatha.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Manmatha

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.985328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.569007Z digest=sha256:e6f08b11995a95704c22c889fba38e8fbe61ee6992930c716f5e17c4ec2f510f

Observation aab5ecd6-5ddd-4812-9a99-a9d0e01cbdea · outbound

This paper cites Tuan Pham, S.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Tuan Pham, S

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.975515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.572295Z digest=sha256:e9392be650450d9a4bbd713d3b572dcc12c7c9f2f456f8d580170b230f6f0e08

Observation 3340cb68-c9a3-433b-ac70-c5a9fc519aa3 · outbound

This paper cites Model cascading: Towards jointly improving efficiency and accuracy of nlp systems.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Model cascading: Towards jointly improving efficiency and accuracy of nlp systems

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.965723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.575573Z digest=sha256:a872c036ca3ac955ef3bded6315c796eeb0bbf2ffb8bf362e14d5b1fdb072960

Observation a84b56e0-c4d7-43ac-9b3e-6fa0effefe23 · outbound

This paper cites Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Investigating acceleration of LLaMA inference by enabling intermediate layer decoding via instruction tuning with ‘lite’

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.955798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.578809Z digest=sha256:82b9e1e83c318e8cfe1b50e211349910ae27fa52e9556c9989990a2f5f18f4f9

Observation eaccb3ce-68cd-4c3f-86a9-c4aa429352f3 · outbound

This paper cites Attention is all you need.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Attention is all you need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.581952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.581952Z digest=sha256:8bf43fee2c02001521864c4584529b0a14ba168d005c07ebab61598f7c16647e

Observation 941f8570-0355-4682-9b53-c23bb9ffce91 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.585245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.585245Z digest=sha256:9443f49f5bad5399140b881fe32144fe5731c1c0aeae39c5090aa6d209758817

Observation eb6434e6-a44b-44f0-8041-5c12645a8b5f · outbound

This paper cites Transformers: State-of-the-art natural language processing, 2020.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Transformers: State-of-the-art natural language processing, 2020

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.940331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.588725Z digest=sha256:797db5c78f6f04635767f7a49818bb6cd9ad67076ca4747b00d0ea173231d8cc

Observation f6407729-53b6-45dd-acbe-16cf625e36cf · outbound

This paper cites Speculative decoding: Lossless speedup of autoregressive translation, 2023.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Speculative decoding: Lossless speedup of autoregressive translation, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.930511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.591822Z digest=sha256:92ea17c490d435815dabfca1cfb32e3d2b87a34de9f58cb3cf7b6aa7238dfd30

Observation 09a028c5-939f-4d78-ba5c-e2473c964c4d · outbound

This paper cites Deebert: Dynamic early exiting for accelerating bert inference.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Deebert: Dynamic early exiting for accelerating bert inference

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.920409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.595373Z digest=sha256:df9678b025ced02204eb36e008f6d8eddaef46c86a97f2f15f645ccdc8dba0ac

Observation 772efeba-ed59-4ef5-a62b-88fe4ebbe676 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.598812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.598812Z digest=sha256:38111c699b9f5108474e5ec4a4f4fc0d1de9baabea68a0f73aebe821a6195d6f

Observation a3788be9-4e4b-4053-ad3c-d74088f4b257 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.909812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.602300Z digest=sha256:26486f37317476ca2d0d18e04f9239d1639011d93aec128ff98631de31c52314

Observation 350e3e94-af68-4af6-9d1a-37eba339cc6f · outbound

This paper cites Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Spider: A Large-Scale Human-Labeled Dataset for Complex and Cross-Domain Semantic Parsing and Text-to-SQL Task

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.605732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.605732Z digest=sha256:1bb9bc9048c2ed5046e95a39d3d4054d9c1cdeb6b82cb2d00218f1893c02a719

Observation e22e5582-7e5d-45ca-aeec-b77e78ccb127 · outbound

This paper cites P Xing, Hao Zhang, Joseph E.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference P Xing, Hao Zhang, Joseph E

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.609457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.609457Z digest=sha256:cd31965e70fa0dd6e5fe44be8489f1941500d131caebed273babe6994f884c90

Observation 334aef28-f2e3-4be0-b144-280af8aee97c · outbound

This paper cites Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-09T13:40:56.612810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T13:40:56.612810Z digest=sha256:fb2f40cde8e4209b8f476b64ecd811db89dd0b9ab1b75c9eb02bb378a8b8bcbe

Observation da1245c6-485f-4294-88a2-171311cd026f · outbound

This paper cites Zhou and et al.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Zhou and et al

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.893641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.616348Z digest=sha256:e4f2c644f9dd2f48bf73586703376a41bc60adb3d4b0e3407c49d28d5e499b2e

Observation e687d854-34be-4b62-9d06-1176b7fdb6f1 · outbound

This paper cites Designing efficient sparse expert models.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Designing efficient sparse expert models

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:40:56.883584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.619675Z digest=sha256:d84adc9c43d3b8a72ae49eb5229d3c4f9b1d0ac03a697da572922d955e27fb93

Observation f72b9031-d6d4-4fbb-9136-bdaa7c543a1b · outbound

This paper cites an unresolved cited work.

M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-09T13:40:56.874916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-09T13:40:56.623496Z digest=sha256:ef4ccadf000da2237fd1298c728e5c1202e40d7d2916eaf4bcf50dc84e6905d5

Pith citing papers

No inbound Pith citation observations are available.