Pith. sign in

Paper Citation Record · LEDGER

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters

As of 9 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2502.07832.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07832 v1

Coverage vector

measured 98 of 98 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.720633Z

measured 98 of 98 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

98 of 98 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved82
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 609a0388-61f5-4d4f-8bab-334ba9f061f7 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.394550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.394550Z digest=sha256:5d029978bab937edcec92914ec20243c3d8e5a3cfd43cf054a870edebd984fb6

Observation f78eda28-a31b-4362-9066-1676b007fe72 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.399688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.399688Z digest=sha256:68946bdc6ff71b49f8d475e91b4f6adb9d32dbafafeaa707adb4493a366a5314

Observation ad89bf29-a54e-4554-993e-c2a38d02cae8 · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.403441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.403441Z digest=sha256:ebb8ec27718c11802f57ead225387dd819a501d82face1a56e965e28978ef1b6

Observation 4c65fb43-6d32-422b-b5f5-f7e98e5c1f8b · outbound

This paper cites Qwen Technical Report.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.407280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.407280Z digest=sha256:2a4643e90bf498eb461ac10b8b5277aa1a23e480bdd8ae28b7c96cdfc430f93e

Observation 5219e199-9660-4a5e-ba13-6eda36712582 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.411395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.411395Z digest=sha256:c36c8620f72712b727af5020e841eba913241515cd666fc4c18182ece2f97ae0

Observation adf91926-3d4f-4a60-a854-f1f9a93f7d7b · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Piqa: Reasoning about physical commonsense in natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.415307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.415307Z digest=sha256:18704fe65bbe08841c18ca2d91623f9f3514fd10b73171362d9e781235a385f5

Observation cc9c7778-8471-48d6-ac99-e2e7b0a8d294 · outbound

This paper cites GPT-NeoX-20B: An Open-Source Autoregressive Language Model.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPT-NeoX-20B: An Open-Source Autoregressive Language Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.419177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.419177Z digest=sha256:3bdf95953161503def38ef7c3d880b442a5787782eb071f690e622e06fdeba0b

Observation 2fa7b9b5-d346-44e6-a89f-9ac1bcf1e561 · outbound

This paper cites Language Models are Few-Shot Learners.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language Models are Few-Shot Learners

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.422737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.422737Z digest=sha256:8883bff4cbe0807da7b33de1db6a37bbdc63fdf91a9a460c353c0b303cd44c05

Observation 33217203-abf2-4e99-9754-d8f3a24df587 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.426468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.426468Z digest=sha256:553d467fd60d3b310a4b6619be85a28973cb2ee2e5296f11105d5a14e0b3b880

Observation 99af52a5-86ed-477c-97da-3b0ec35d1cdc · outbound

This paper cites Code alpaca: An instruction-following llama model for code generation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Code alpaca: An instruction-following llama model for code generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.429889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.429889Z digest=sha256:254ebe66995602a846045d08669b9e3b753d502eda63bf95c283b68387956f05

Observation ac8a45f6-28b3-4949-b2bd-0700b1623bde · outbound

This paper cites Learning to maximize mutual information for chain-of-thought distillation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learning to maximize mutual information for chain-of-thought distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.433149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.433149Z digest=sha256:322802c9702f0cd0cf7e76ec45856b5848d78d956704486702bb527858a3e64b

Observation e4d976e9-10fa-47c7-b95f-0791502939b2 · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.436326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.436326Z digest=sha256:6cd14281f0ad5579bf26cede3faab5539ff26cedc80f8f0d8ca54804a954423f

Observation 0dda9464-5803-46e1-828d-e2c8405ddf5c · outbound

This paper cites D ialog S um: A real-life scenario dialogue summarization dataset.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters D ialog S um: A real-life scenario dialogue summarization dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.439774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.439774Z digest=sha256:80274b03fb83e33b94ddda8356068d110005c7fac582584a1b0938cab0be5d48

Observation dca0e07c-86e1-4425-bf77-755de8711dbe · outbound

This paper cites Palm: Scaling language modeling with pathways.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Palm: Scaling language modeling with pathways

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.442950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.442950Z digest=sha256:275a86d73965da541d95f37ad114ba29e19c93d6295788603193a3ba913771b1

Observation af5fd9e9-6e76-4b30-9820-56d76ac01d54 · outbound

This paper cites Boolq: Exploring the surprising difficulty of natural yes/no questions.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Boolq: Exploring the surprising difficulty of natural yes/no questions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.445935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.445935Z digest=sha256:394f8c4f63bd471c76f4d1c7b9fd962f5b82fb4dd25ef0e6e78cda1434c43527

Observation 76976001-ca3e-4cc4-8e76-993290329eff · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.448983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.448983Z digest=sha256:fa159dd2579942ffbeb3c0a0aadc307f245a483f08a634ffca261d60f20d8501

Observation 384bbdd8-c598-4cc6-8732-558b9d69691f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Verifiers to Solve Math Word Problems

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.452482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.452482Z digest=sha256:f1765d9fb44a5c2224b5acb8ee0da5eb0bd2dbb902ecd9399a799fb78d568f4f

Observation 55dcea1e-fcb7-4a65-85bc-ba36521e10a9 · outbound

This paper cites Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.456036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.456036Z digest=sha256:9c501aaa340141a408f991f659d6d0dcdf6f10f1d295e9139b13c5e51f60e41a

Observation 5144d1d7-b9ea-43d0-b3b9-ac95b4ea587f · outbound

This paper cites Mutual: A dataset for multi-turn dialogue reasoning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mutual: A dataset for multi-turn dialogue reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.459132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.459132Z digest=sha256:f8c73f9456d8292749837de1cbb542cbbdca2442a2f8ae2c10109245e1831f31

Observation f4f976b7-223a-4a05-98af-bec7aced9cd8 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.462157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.462157Z digest=sha256:514c057eccefdc3ef38bc74e776a3b5f7e4b8dbc2ed64db05482ad9279be5245

Observation cd351f96-3ef2-4134-8c23-5cb350db3938 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.465407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.465407Z digest=sha256:06517bedce52862765da6ba15ca87b679cddd93440c9e4ea72240d1f0c1b8ede

Observation d2947059-ec14-420b-b640-69bc0a070c9d · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.468388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.468388Z digest=sha256:ed6c12be51b2e23a106a3690d22f3881fcfdb4d20934c9096cba3083f9c988e4

Observation f04ece8a-781a-4cf6-981a-a8cf6b514c72 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qlora: Efficient finetuning of quantized llms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.471354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.471354Z digest=sha256:7b70b4ad0dc602210338d3fc16aef9ebf4ba8dcecbf58ccecf5cf47ea0778449

Observation 8d69a650-a8bd-4bfe-bada-f1598e01d261 · outbound

This paper cites Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.474327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.474327Z digest=sha256:d81eefc8779564f20215f084ecf61ff33faa2d6a9469fd5e6aef9c3ee585f880

Observation 9bb8c23d-e8c0-4e2b-bae4-dbe1d46a9d9a · outbound

This paper cites Blockwise compression of transformer-based models without retraining.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Blockwise compression of transformer-based models without retraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.477757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.477757Z digest=sha256:e3140db10490497947e66ef5adb1e887a9123c3d514ecac5f733a95cf0224e4f

Observation 6d9b0a07-1127-4b4f-8b2b-b5fab7e721f3 · outbound

This paper cites The Llama 3 Herd of Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.480714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.480714Z digest=sha256:6fec887794aba79c572edf3a3f7021e4c7c82c6503652b3df5a703ae8261be18

Observation d1547bb2-eb02-4cf9-ad61-c92ecfa99937 · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.483877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.483877Z digest=sha256:7a499827d29dadec029408b254632becc256a41ffa4bb4e0110de7656f2d0723

Observation 9373bec1-19a2-4e7b-855b-716f523cf8a0 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.486906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.486906Z digest=sha256:ceafcbc5b989bbea2afe969b16bc3e283fad17aa6f191787d894c1622b8c4abe

Observation d554e6de-e130-4979-a77b-ffaab0f0b5f7 · outbound

This paper cites Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.489949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.489949Z digest=sha256:ca3fb9cb94261c15f76fd57b64546cd6496bd019bc29e85ce89ad70e625b4017

Observation f77d0b6a-ef8e-45b7-aa1e-c0768c072392 · outbound

This paper cites A framework for few-shot language model evaluation, 07 2024.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A framework for few-shot language model evaluation, 07 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.492841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.492841Z digest=sha256:b5cd631ac7b562f0aee2f469812036893fba90898d6d6f4b72c0835ca41691d8

Observation a2a0077f-9906-4d40-a7c2-332cf89e403b · outbound

This paper cites The Unreasonable Ineffectiveness of the Deeper Layers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Unreasonable Ineffectiveness of the Deeper Layers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.495709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.495709Z digest=sha256:7b6c7561d4531878e08a7d9579e1469962d7cae2c5774ed89a7f9d0c43611748

Observation 990270c8-cd96-4223-a541-89660cfc90ba · outbound

This paper cites Textbooks Are All You Need.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Textbooks Are All You Need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.498897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.498897Z digest=sha256:86ffaf4b59274b15c68710f919784f7f059e304267cc20c37dc4aeff586cd4b5

Observation 126127a4-5cba-4546-9584-de4734c56e1e · outbound

This paper cites Compressing pre-trained language models using progressive low rank decomposition.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Compressing pre-trained language models using progressive low rank decomposition

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.502449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.502449Z digest=sha256:6326eb1127b9d9c61d073fc7c0ab4737e36951c7383bb22378d8db8726668763

Observation 1bdf172f-4dc9-4564-81cb-d56f9a209242 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Distilling the Knowledge in a Neural Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.505557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.505557Z digest=sha256:4eddbab6f4c49c4e112a489c37743bc48055b7d0af3a4c681af9c2f88c134954

Observation b740735b-53ed-46ac-8c77-911256352e1a · outbound

This paper cites Training Compute-Optimal Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Compute-Optimal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.508785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.508785Z digest=sha256:3a9c3db948c9a50be86cc668546c7891100a642b1be13bc1f2c5de6b5e94bb4c

Observation 24872e0f-5eed-430f-a35c-01ad60f58d30 · outbound

This paper cites Language model compression with weighted low-rank factorization.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language model compression with weighted low-rank factorization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.512142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.512142Z digest=sha256:05f0a3f203d0adf5df047d915eb45796eb89deda243f264945d39229adeb32f4

Observation 2eae92fb-3812-48f7-b5b7-4b9ef31c70f8 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoRA: Low-Rank Adaptation of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.515423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.515423Z digest=sha256:d11087a2515fbe681dcdd790f1d538d875344752430a295cef2e06f8215f76e9

Observation ec55e8c8-09c2-491b-bdac-9477daf7f5fe · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.518499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.518499Z digest=sha256:42e43b63abefc66aa23b062e488586a1e4916cd09e55fc7be6802fb81a9f907f

Observation d397db10-374b-451a-a7ef-ff47182143d8 · outbound

This paper cites safetensors.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters safetensors

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.521757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.521757Z digest=sha256:f23bf5f77fd59edc8890c2f50dd5fe45ab9fc01363b8ffb44fae2ea027d648bb

Observation 9bebc7d0-9c82-473a-8ad4-c3ac6329add6 · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.524776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.524776Z digest=sha256:4c8ad618f6eb559e40a3a5f6e991b2b6f45efc1188121b23152630da42d6a086

Observation dbea67cc-d6dc-4143-8b6c-4fdf9f4d9c7c · outbound

This paper cites Mistral 7B.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mistral 7B

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.528254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.528254Z digest=sha256:c4c42121eefae0d8e450ff76d2a836195a5bb5f1ca2eef432c2e334fdbb1b95f

Observation 1146dde1-570b-456e-b301-574fa10fb48e · outbound

This paper cites Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:01.008160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.531798Z digest=sha256:92a21791da4dcfefbaac3513fcc162b3f70413a246c91c446cc256cb567fbbe5

Observation d05d8b5b-34a8-4366-8d4e-c58a586a8720 · outbound

This paper cites Weld, and Luke Zettlemoyer.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Weld, and Luke Zettlemoyer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.535267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.535267Z digest=sha256:9c74ed9c2da2ff312a5d5f0df938519e067df16ce9117bd29bc738e323c7c3b4

Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · outbound

This paper cites GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.538291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.538291Z digest=sha256:5edd308ebbc0f0bc83ac28cf152d548b650973603ec658269afa59749fc41281

Observation df2d3abd-ca64-4747-977d-7f5f91b1e126 · outbound

This paper cites arxiv-math-instruct-50.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters arxiv-math-instruct-50

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.569680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.541523Z digest=sha256:10822dc2adb6f2b57d8487530f4676cfa0de960814b7f2a263d3bf3701131201

Observation 82fa24fe-d6f5-4e80-a3b1-bf4e25e3196a · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters SqueezeLLM: Dense-and-Sparse Quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.544467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.544467Z digest=sha256:e3a4b66ada89b80c81677b46c565bf15b45fad882d5a26b1a5fda300070ac737

Observation 1dce0f8e-e1f0-4d76-84ba-eaf82b5e4aa3 · outbound

This paper cites Full Stack Optimization of Transformer Inference: a Survey.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Full Stack Optimization of Transformer Inference: a Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.547913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.547913Z digest=sha256:8d583487d4e790bfa5cda2b137a572143e10f713f04207a99910622251d7c96d

Observation bbe9f86d-16ac-4127-8751-b4fb6bcc56f5 · outbound

This paper cites Speculative decoding with big little decoder.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Speculative decoding with big little decoder

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.559923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.551373Z digest=sha256:81ed85d8a67654162079d59fb4d5d558d25cfd1906ddcb069fd0eb9b6cb8241b

Observation 99319401-96ec-478c-95e5-3543a552728b · outbound

This paper cites Reformer: The Efficient Transformer.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Reformer: The Efficient Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.554454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.554454Z digest=sha256:198737f894f4472b19bbbb75b82a5ea64599c0b86f4cbab89aeab9afd5935bc5

Observation 5605fd54-6203-4d08-841f-2d13198d77bf · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters o pf, Yannic Kilcher, Dimitri von R \

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.550621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.557832Z digest=sha256:6182e7068532d9f0f0a58a76decde7f7487bbef058ab20dc77716eadef7409b3

Observation 448f8bc4-5c02-4654-a90c-a552d35cd2f8 · outbound

This paper cites LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.560727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.560727Z digest=sha256:e54d9be900940f766f0c5964bb73e7bb9308663c81c56e206dc257d5c14bd8e7

Observation 1ce2381e-10f6-46c4-8bc8-c2e94232da5f · outbound

This paper cites Losparse: Structured compression of large language models based on low-rank and sparse approximation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Losparse: Structured compression of large language models based on low-rank and sparse approximation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.541682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.564027Z digest=sha256:0c025a8583275e5db48d8a565a2200c8312826ced8c6052452507bd912e6d724

Observation 08840d56-194c-45a0-9a03-329bef4349d9 · outbound

This paper cites The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:00.955687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.567152Z digest=sha256:c191236f042e4734cfcda918eee2d60f11f9ce9a317bc063f83489b7b0ae33ab

Observation 9bbb9f5b-f553-4d78-9461-5702cce38c02 · outbound

This paper cites LLM-QAT: Data-Free Quantization Aware Training for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LLM-QAT: Data-Free Quantization Aware Training for Large Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.570594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.570594Z digest=sha256:142a8e942d02bc56c34572f0a945b6920ad7b51a68c4274a11f96c6c384a8734

Observation a1aac243-d292-44e6-8a38-1538e7bab617 · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.574242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.574242Z digest=sha256:c797fd4efb5e6009974192f981fc4e275611da958006e97245eb92b4892338a3

Observation c7c1be55-0710-45b8-8297-1609f64d69ef · outbound

This paper cites Deja vu: Contextual sparsity for efficient llms at inference time.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deja vu: Contextual sparsity for efficient llms at inference time

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.532044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.577593Z digest=sha256:58de1eb940ba24ecd8374f049b4bbc95fd3f243e07b4fc4ef0728f1051c9cd5d

Observation d63f4650-1139-4eed-848c-01ff37af7fe7 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The flan collection: Designing data and methods for effective instruction tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.580675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.580675Z digest=sha256:0266c6431f0bd8850166bee322cfc1b47bd83944019f3d3e1d86054f7b36e1b3

Observation f1f98760-770e-4885-b5f7-eb28ddd2e9b4 · outbound

This paper cites Fineweb-edu, May 2024.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Fineweb-edu, May 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.518017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.583970Z digest=sha256:48f6c376bd376ba22d6f88c3c9adee6488082cd8eeffcc55bdd64c018338590e

Observation 029c6bc1-782e-44a0-aa33-f43bad5966ed · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llm-pruner: On the structural pruning of large language models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.587134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.587134Z digest=sha256:eb67924cbb55aadff4d501d58fc654532e881fd896a0966e982eb3c786d6ee1a

Observation 4357665d-8a38-4c5f-ac05-eba6b630ff83 · outbound

This paper cites Peft: State-of-the-art parameter-efficient fine-tuning methods.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Peft: State-of-the-art parameter-efficient fine-tuning methods

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.590146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.590146Z digest=sha256:914a7363176116fb6807489f5d5fa793d2059c9a9a6beab3894b597b8a2375ab

Observation b1b78aea-63f9-44a5-ac99-833b591b03f4 · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.593433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.593433Z digest=sha256:650d9802223ec1390fd0b8ba200728e4b412099d46354e31eaea295e2c01fff7

Observation 8c46f36c-57ce-453b-8ee9-57271195c5ea · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.598069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.598069Z digest=sha256:b64933e0d4af4f93b1798c733d7a782342f621e4b53081c91220b8ed60b5e78e

Observation 2c0e37d9-602e-4b4e-b89a-2eae112783c5 · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.498518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.601574Z digest=sha256:232c25100432f18bd2fbbc5c51dbc1d524043d0120d96861f737bc036242e6c7

Observation cd3188e7-a885-4cf6-adbb-99ab332cd029 · outbound

This paper cites The lambada dataset, Aug 2016.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The lambada dataset, Aug 2016

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.604742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.604742Z digest=sha256:b6da51db88e92043521056f782baea772476dc0fa76d9c1705325634e3672247

Observation 73e2852b-3fd7-438a-bbd9-86f2ff6cf807 · outbound

This paper cites Hovy, Pamela Forner, \'A lvaro Rodrigo, Richard F.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Hovy, Pamela Forner, \'A lvaro Rodrigo, Richard F

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.482413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.608043Z digest=sha256:6a6c3f7700dad445eefb51ba0312f50f878a2f688fd025fdb2c2e34beaad8aa8

Observation 94a895b8-0d03-49ae-a33b-828a0dc37633 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.610889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.610889Z digest=sha256:706b8cfd1ca50f97f8af19eeb306e0af32f817bd99d74cc29ca69aa80a30188a

Observation 9ae97c1d-32cb-4a28-bd85-2fe2b0b93c53 · outbound

This paper cites Instruction Tuning with GPT-4.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Instruction Tuning with GPT-4

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.614144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.614144Z digest=sha256:58ba708a14bc305d29f3e1f29e799db9962c08fa250565ad8eccf8855f22cc0f

Observation 9f48c874-de2e-4c5d-af16-d8b5f722c1f3 · outbound

This paper cites Efficiently scaling transformer inference.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Efficiently scaling transformer inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.617414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.617414Z digest=sha256:dbba4290b15fca80473e64186b9e3b6280d6c77afbc26e764a2175e4d085e8fa

Observation 05bb520b-de63-45e1-ac97-2cefe282825b · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.620603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.620603Z digest=sha256:cf13d89884b9d540dc83f67574a22bf9903f67b4fbc3748cffeca3e0e542b5f0

Observation 7d118554-e8be-428f-b59f-ac40539283d7 · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.623850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.623850Z digest=sha256:7d8fa359f18d1cdd7f1764b221efef4c5e53c7b06ab073df4cef35ed2faf6ea2

Observation 2bc13ed9-f377-440b-89eb-7d0e1d2795b8 · outbound

This paper cites S-LoRA: Serving Thousands of Concurrent LoRA Adapters.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters S-LoRA: Serving Thousands of Concurrent LoRA Adapters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.627303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.627303Z digest=sha256:02a9f68b41219445706ea2140836a44b65153b3ccd1cd4390c8bc5b46e951f20

Observation 4429da89-7e1a-4f88-8de7-731eb9f950dc · outbound

This paper cites Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.630570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.630570Z digest=sha256:aa8f62596c7ec6ca8d095715a9740406966d7b60b28da070ea0c979f404746c3

Observation 79f7021f-7e3c-4f95-ad0f-c03563279973 · outbound

This paper cites Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.467507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.634440Z digest=sha256:6c197441154933d31a8209f4a83404c86565c5dffe0df78ca457893915ed6253

Observation d0863e9f-5c08-4595-a736-0465ac3e2d42 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A Simple and Effective Pruning Approach for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.638205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.638205Z digest=sha256:99c17a601860ed9fb45a3bba5d84ca8e6345169c0a07a43f0fcd0daf291dbc6a

Observation 83636a88-52f1-4891-b66d-98e554e2fb56 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.641561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.641561Z digest=sha256:023b0ed47fc0692acfc2d3cfe99b6f64c453daad94d2846bbc133ea8ecc3b623

Observation aa25504b-1c32-4c06-8ab1-ac8be1fea274 · outbound

This paper cites KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.645377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.645377Z digest=sha256:fb19067c533391f088e7552e42765b55d7c915ab718dca0ffc64c6d1d6bb6dbc

Observation ba86a792-cd7d-4434-8f78-c84ca18c9055 · outbound

This paper cites C ommonsense QA : A question answering challenge targeting commonsense knowledge.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters C ommonsense QA : A question answering challenge targeting commonsense knowledge

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.648906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.648906Z digest=sha256:2b29063e1ebb7a7caedff0899d8bee48a13979c4fdd247bce135703389a1cdd3

Observation 4414b662-01c4-436b-a618-9f0b6af1aae3 · outbound

This paper cites Multi-Domain Neural Machine Translation.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-08T13:44:00.837719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.652312Z digest=sha256:27a256626c8ca0da6fb6ec2641f31d68fc722682db3f7c114081f0d51703a3cf

Observation 6025f9a1-8fe8-4a9d-88e3-8abc8d210e5f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Gemini: A Family of Highly Capable Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.655703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.655703Z digest=sha256:3ab904f5548f03420f3b7a14060f959e08d70732c2f9d518d2424aea8ebaa2d7

Observation 878a148a-0bcb-4fc8-86d7-4e755cb41522 · outbound

This paper cites Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.659142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.659142Z digest=sha256:7de35e6b03ce4d04934ce25e68627b41adfd5bc819ecc9cbd10e4f0b6827191e

Observation f3aa82dc-5707-4ac0-a5d0-6748380455f8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.662881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.662881Z digest=sha256:b34868c722af3226e2eef178072eecb63928aa602f206daa9d5fe0b0a761931c

Observation a189eff1-a16f-4fbc-9d37-3025366e41d4 · outbound

This paper cites Smith, Iz Beltagy, and Hannaneh Hajishirzi.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smith, Iz Beltagy, and Hannaneh Hajishirzi

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.666585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.666585Z digest=sha256:a8eba315477b32e3da7659cc5b52fc3960b8d475874780f02002069be84dd9db

Observation 6090a7a2-414f-4013-9fd5-43057ea2c880 · outbound

This paper cites Liu, and Matt Gardner.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Liu, and Matt Gardner

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.452340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.669764Z digest=sha256:529324efaa5d3aa87e82c39cb5283455aa9a434b5c4ef943af5541aa7e0553ee

Observation 72d3a5b1-ca63-4a16-bac2-3fb80da30ac2 · outbound

This paper cites Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.672927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.672927Z digest=sha256:4f958aff8541a053783274bc28f690634bd8f1c530ff478566c1526596ecbcf2

Observation 197c667f-0f91-4cb0-9932-05469fef679c · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.676353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.676353Z digest=sha256:81e6b25f7f07c4e0792e8e21f2a5c203b2402d54d07c2775391ae7c20bb02fa9

Observation 398d67a9-e006-4337-9adf-55117488af82 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.680017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.680017Z digest=sha256:c0067a5335774c0d0df5f143bc3dee5273f2e064e1dc7f6592f317a253871981

Observation 006da0f2-6bb7-4c61-bd0d-254ff8a5779e · outbound

This paper cites Wizard LM : Empowering large pre-trained language models to follow complex instructions.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Wizard LM : Empowering large pre-trained language models to follow complex instructions

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.683256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.683256Z digest=sha256:fd7ccc2ceceabfc3dfc46db0b8ac880541599bc9e8bfe7298900378a91e6440d

Observation a40591c7-eb13-4dae-9979-8ad86ed6550b · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.432369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.686408Z digest=sha256:960f9aa6bcc84d96355139452daf8bcc7d2861713f747c2ad565625138437211

Observation d552fbc1-bcdd-4087-80e6-6d7e01384118 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters TinyLlama: An Open-Source Small Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.689555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.689555Z digest=sha256:d12d928fe8a0e7351248c713e55c18ab04322684b99cb4864886af35c95a18ae

Observation e2ada6c8-054c-4cf5-b556-dffa5ae0fd7f · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.692882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.692882Z digest=sha256:621a13068fc3550516f41cd17c920a9de54aff4bf5badb14e29d1318a4ab8dbb

Observation 6a0cda3f-0f67-4f00-b8db-30140497589f · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters OPT: Open Pre-trained Transformer Language Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.696184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.696184Z digest=sha256:3fc533bfaf35dab240af28a984dda7e43d86950948e2e2d4f384c5c913722b45

Observation 788b02db-186b-4043-aaa4-f3b04a13b843 · outbound

This paper cites Adaptive-precision framework for sgd using deep q-learning.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Adaptive-precision framework for sgd using deep q-learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.423105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.699834Z digest=sha256:c3b45347a48452b013fd213a19b6899ed4abe37b873d309e4a9fc0478f9e924e

Observation f457a825-98a0-48cd-883c-08043957a28e · outbound

This paper cites H2o: Heavy-hitter oracle for efficient generative inference of large language models.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters H2o: Heavy-hitter oracle for efficient generative inference of large language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:44:01.413628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T13:44:00.702933Z digest=sha256:29ea6c8c59ada66501c6a492f589ac16634c36744fdde14fe6b476b61d1478cd

Observation 77948001-69c7-4cd3-8931-ccdf9ca04e35 · outbound

This paper cites Lima: Less is more for alignment.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Lima: Less is more for alignment

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.705824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.705824Z digest=sha256:e7893a8b21a63b22a957df071dd4ed00752c2af1ba9be635e518342dcf633106

Observation 1e7fd59a-1c84-4197-b6b6-2ef9c4a72fd9 · outbound

This paper cites write newline.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.708804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.708804Z digest=sha256:9363c575e9e0c969be088c276d28e2a6aa1a475bd53fea3d67c16c0689ec88c0

Observation f32fecc0-8bc8-4ca4-ab83-82106412c486 · outbound

This paper cites @esa (Ref.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters @esa (Ref

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.712425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.712425Z digest=sha256:13e588c8421898e69509900eb823d63d66cc40eba48ac9703eabb7ac4e5df721

Observation 669dfaa8-092f-47d8-970c-c4cf8d21571f · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.717308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.717308Z digest=sha256:43df91f5153104254117a37080b6ebdbae672e851d10b4088ef0bdc5ca5f7cde

Observation 555edb9b-f77f-4b11-a6a6-97964e650b49 · outbound

This paper cites an unresolved cited work.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.720633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.720633Z digest=sha256:bd137eff394de66a5fc3e289de34c8877623008379dc65437b62e80dbca54f90

Pith citing papers

No inbound Pith citation observations are available.