Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.720633Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 98 of 98 outbound references and 0 inbound Pith citation observations for arXiv:2502.07832.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:44:00.720633Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
98 of 98 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 609a0388-61f5-4d4f-8bab-334ba9f061f7 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f78eda28-a31b-4362-9066-1676b007fe72 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad89bf29-a54e-4554-993e-c2a38d02cae8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mathqa: Towards interpretable math word problem solving with operation-based formalisms, 2019
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c65fb43-6d32-422b-b5f5-f7e98e5c1f8b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qwen Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5219e199-9660-4a5e-ba13-6eda36712582 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Pythia: A suite for analyzing large language models across training and scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf91926-3d4f-4a60-a854-f1f9a93f7d7b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Piqa: Reasoning about physical commonsense in natural language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9c7778-8471-48d6-ac99-e2e7b0a8d294 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa7b9b5-d346-44e6-a89f-9ac1bcf1e561 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language Models are Few-Shot Learners
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33217203-abf2-4e99-9754-d8f3a24df587 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99af52a5-86ed-477c-97da-3b0ec35d1cdc · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Code alpaca: An instruction-following llama model for code generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8a45f6-28b3-4949-b2bd-0700b1623bde · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learning to maximize mutual information for chain-of-thought distillation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d976e9-10fa-47c7-b95f-0791502939b2 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dda9464-5803-46e1-828d-e2c8405ddf5c · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters D ialog S um: A real-life scenario dialogue summarization dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca0e07c-86e1-4425-bf77-755de8711dbe · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Palm: Scaling language modeling with pathways
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5fd9e9-6e76-4b30-9820-56d76ac01d54 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Boolq: Exploring the surprising difficulty of natural yes/no questions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76976001-ca3e-4cc4-8e76-993290329eff · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 384bbdd8-c598-4cc6-8732-558b9d69691f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Verifiers to Solve Math Word Problems
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55dcea1e-fcb7-4a65-85bc-ba36521e10a9 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Free dolly: Introducing the world's first truly open instruction-tuned llm, 2023
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5144d1d7-b9ea-43d0-b3b9-ac95b4ea587f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mutual: A dataset for multi-turn dialogue reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4f976b7-223a-4a05-98af-bec7aced9cd8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd351f96-3ef2-4134-8c23-5cb350db3938 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2947059-ec14-420b-b640-69bc0a070c9d · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04ece8a-781a-4cf6-981a-a8cf6b514c72 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Qlora: Efficient finetuning of quantized llms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d69a650-a8bd-4bfe-bada-f1598e01d261 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb8c23d-e8c0-4e2b-bae4-dbe1d46a9d9a · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Blockwise compression of transformer-based models without retraining
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d9b0a07-1127-4b4f-8b2b-b5fab7e721f3 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Llama 3 Herd of Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1547bb2-eb02-4cf9-ad61-c92ecfa99937 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sparsegpt: Massive language models can be accurately pruned in one-shot
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9373bec1-19a2-4e7b-855b-716f523cf8a0 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d554e6de-e130-4979-a77b-ffaab0f0b5f7 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Learn-to-share: A hardware-friendly transfer learning framework exploiting computation and parameter sharing
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f77d0b6a-ef8e-45b7-aa1e-c0768c072392 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A framework for few-shot language model evaluation, 07 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a0077f-9906-4d40-a7c2-332cf89e403b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Unreasonable Ineffectiveness of the Deeper Layers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990270c8-cd96-4223-a541-89660cfc90ba · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Textbooks Are All You Need
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126127a4-5cba-4546-9584-de4734c56e1e · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Compressing pre-trained language models using progressive low rank decomposition
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bdf172f-4dc9-4564-81cb-d56f9a209242 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Distilling the Knowledge in a Neural Network
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b740735b-53ed-46ac-8c77-911256352e1a · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Training Compute-Optimal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24872e0f-5eed-430f-a35c-01ad60f58d30 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Language model compression with weighted low-rank factorization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eae92fb-3812-48f7-b5b7-4b9ef31c70f8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoRA: Low-Rank Adaptation of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec55e8c8-09c2-491b-bdac-9477daf7f5fe · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d397db10-374b-451a-a7ef-ff47182143d8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters safetensors
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bebc7d0-9c82-473a-8ad4-c3ac6329add6 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbea67cc-d6dc-4143-8b6c-4fdf9f4d9c7c · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Mistral 7B
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1146dde1-570b-456e-b301-574fa10fb48e · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation with Word-Level Adaptive Layer-wise Domain Mixing
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d05d8b5b-34a8-4366-8d4e-c58a586a8720 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Weld, and Luke Zettlemoyer
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0872bdf5-f05e-4ce2-b24d-f984a37244e5 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df2d3abd-ca64-4747-977d-7f5f91b1e126 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters arxiv-math-instruct-50
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82fa24fe-d6f5-4e80-a3b1-bf4e25e3196a · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters SqueezeLLM: Dense-and-Sparse Quantization
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dce0f8e-e1f0-4d76-84ba-eaf82b5e4aa3 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Full Stack Optimization of Transformer Inference: a Survey
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe9f86d-16ac-4127-8751-b4fb6bcc56f5 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Speculative decoding with big little decoder
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 99319401-96ec-478c-95e5-3543a552728b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Reformer: The Efficient Transformer
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5605fd54-6203-4d08-841f-2d13198d77bf · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters o pf, Yannic Kilcher, Dimitri von R \
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 448f8bc4-5c02-4654-a90c-a552d35cd2f8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ce2381e-10f6-46c4-8bc8-c2e94232da5f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Losparse: Structured compression of large language models based on low-rank and sparse approximation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 08840d56-194c-45a0-9a03-329bef4349d9 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The Microsoft Toolkit of Multi-Task Deep Neural Networks for Natural Language Understanding
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9bbb9f5b-f553-4d78-9461-5702cce38c02 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1aac243-d292-44e6-8a38-1538e7bab617 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7c1be55-0710-45b8-8297-1609f64d69ef · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Deja vu: Contextual sparsity for efficient llms at inference time
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d63f4650-1139-4eed-848c-01ff37af7fe7 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The flan collection: Designing data and methods for effective instruction tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f98760-770e-4885-b5f7-eb28ddd2e9b4 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Fineweb-edu, May 2024
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 029c6bc1-782e-44a0-aa33-f43bad5966ed · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llm-pruner: On the structural pruning of large language models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4357665d-8a38-4c5f-ac05-eba6b630ff83 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Peft: State-of-the-art parameter-efficient fine-tuning methods
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b78aea-63f9-44a5-ac99-833b591b03f4 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c46f36c-57ce-453b-8ee9-57271195c5ea · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c0e37d9-602e-4b4e-b89a-2eae112783c5 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cd3188e7-a885-4cf6-adbb-99ab332cd029 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The lambada dataset, Aug 2016
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73e2852b-3fd7-438a-bbd9-86f2ff6cf807 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Hovy, Pamela Forner, \'A lvaro Rodrigo, Richard F
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94a895b8-0d03-49ae-a33b-828a0dc37633 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae97c1d-32cb-4a28-bd85-2fe2b0b93c53 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Instruction Tuning with GPT-4
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f48c874-de2e-4c5d-af16-d8b5f722c1f3 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Efficiently scaling transformer inference
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05bb520b-de63-45e1-ac97-2cefe282825b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d118554-e8be-428f-b59f-ac40539283d7 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bc13ed9-f377-440b-89eb-7d0e1d2795b8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters S-LoRA: Serving Thousands of Concurrent LoRA Adapters
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4429da89-7e1a-4f88-8de7-731eb9f950dc · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79f7021f-7e3c-4f95-ad0f-c03563279973 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, Agnieszka Kluska, Aitor Lewkowycz, Akshat Agarwal, Alethea Power, Alex Ray, Alex Warstadt, Alexander W
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d0863e9f-5c08-4595-a736-0465ac3e2d42 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters A Simple and Effective Pruning Approach for Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83636a88-52f1-4891-b66d-98e554e2fb56 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa25504b-1c32-4c06-8ab1-ac8be1fea274 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters KroneckerBERT: Learning Kronecker Decomposition for Pre-trained Language Models via Knowledge Distillation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba86a792-cd7d-4434-8f78-c84ca18c9055 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters C ommonsense QA : A question answering challenge targeting commonsense knowledge
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4414b662-01c4-436b-a618-9f0b6af1aae3 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Multi-Domain Neural Machine Translation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6025f9a1-8fe8-4a9d-88e3-8abc8d210e5f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Gemini: A Family of Highly Capable Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 878a148a-0bcb-4fc8-86d7-4e755cb41522 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Baby Llama: knowledge distillation from an ensemble of teachers trained on a small dataset with no performance penalty
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3aa82dc-5707-4ac0-a5d0-6748380455f8 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a189eff1-a16f-4fbc-9d37-3025366e41d4 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smith, Iz Beltagy, and Hannaneh Hajishirzi
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6090a7a2-414f-4013-9fd5-43057ea2c880 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Liu, and Matt Gardner
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72d3a5b1-ca63-4a16-bac2-3fb80da30ac2 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Flash-LLM: Enabling Cost-Effective and Highly-Efficient Large Generative Model Inference with Unstructured Sparsity
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 197c667f-0f91-4cb0-9932-05469fef679c · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398d67a9-e006-4337-9adf-55117488af82 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006da0f2-6bb7-4c61-bd0d-254ff8a5779e · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Wizard LM : Empowering large pre-trained language models to follow complex instructions
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a40591c7-eb13-4dae-9979-8ad86ed6550b · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d552fbc1-bcdd-4087-80e6-6d7e01384118 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters TinyLlama: An Open-Source Small Language Model
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2ada6c8-054c-4cf5-b556-dffa5ae0fd7f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0cda3f-0f67-4f00-b8db-30140497589f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters OPT: Open Pre-trained Transformer Language Models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788b02db-186b-4043-aaa4-f3b04a13b843 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Adaptive-precision framework for sgd using deep q-learning
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f457a825-98a0-48cd-883c-08043957a28e · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters H2o: Heavy-hitter oracle for efficient generative inference of large language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77948001-69c7-4cd3-8931-ccdf9ca04e35 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Lima: Less is more for alignment
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e7fd59a-1c84-4197-b6b6-2ef9c4a72fd9 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters write newline
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f32fecc0-8bc8-4ca4-ab83-82106412c486 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters @esa (Ref
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669dfaa8-092f-47d8-970c-c4cf8d21571f · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555edb9b-f77f-4b11-a6a6-97964e650b49 · outbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Unresolved cited work
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.