REVIEW 12 cited by
In-context Learning Distillation: Transferring Few-shot Learning Ability of Pre-trained Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Given the success with in-context learning of large pre-trained language models, we introduce in-context learning distillation to transfer in-context few-shot learning ability from large models to smaller models. We propose to combine in-context learning objectives with language modeling objectives to distill both the ability to read in-context examples and task knowledge to the smaller models. We perform in-context learning distillation under two different few-shot learning paradigms: Meta In-context Tuning (Meta-ICT) and Multitask In-context Tuning (Multitask-ICT). Multitask-ICT performs better on multitask few-shot learning but also requires more computation than Meta-ICT. Our method shows consistent improvements for both Meta-ICT and Multitask-ICT on two benchmarks: LAMA and CrossFit. Our extensive experiments and analysis reveal that in-context learning objectives and language modeling objectives are complementary under the Multitask-ICT paradigm. In-context learning objectives achieve the best performance when combined with language modeling objectives.
Forward citations
Cited by 12 Pith papers
-
Sample-Efficient Learning from Agent Experience
Agents can consolidate in-context trial-and-error learning into their weights by distilling one-step teacher decisions at recorded histories, needing no additional environment interaction.
-
FTP: A Fine-grained Token-wise Pruner for Large Language Models via Token Routing
A token-wise pruner with a learned router and a genetic-algorithm sparsity scheduler claims near-lossless LLM inference at 22-40% token sparsity.
-
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
PEFT adapters trained on high-resource summarization domains can improve Llama-3-8B's summaries on unseen domains, but the reported gains are weakened by test-set selection and missing significance tests.
-
AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes
AgentDistill distills agent capabilities without any training by having a teacher generate reusable MCP tool boxes that small-model students invoke at inference time.
-
Extracting Interpretable Task-Specific Circuits from Large Language Models for Faster Inference
A greedy, ablation-based pruning method extracts a standalone task-specific subnetwork from GPT-2 Small, reducing parameters by up to 82.77% while keeping accuracy on three synthetic single-token tasks.
-
Exploring the Limits of Model Compression in LLMs: A Knowledge Distillation Study on QA Tasks
Distilled students at 43% to 50% of teacher size keep over 90% of teacher Exact Match on SQuAD and MLQA, though one-shot gains reverse on SQuAD test for Pythia.
-
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
RLRC combines 90% structured pruning, supervised fine-tuning, PPO reinforcement learning, and optional 4-bit quantization to recover OpenVLA's success rate on ManiSkill while cutting memory up to 8x and boosting throu...
-
SWSC: Shared Weight for Similar Channel in LLM
SWSC combines channel K-means clustering with an SVD low-rank error correction to compress LLM weights, and reports lower perplexity than RTN quantization on Llama-2-7B Q and K projections at 2 to 3 average bits.
-
LSAQ: Layer-Specific Adaptive Quantization for Large Language Model Deployment
LSAQ assigns higher quantization precision to layers deemed important by the overlap of top-k input and output token sets, and reports small accuracy and perplexity gains over a cosine-similarity baseline.
-
Scaling New Frontiers: Insights into Large Recommendation Models
Deep HSTU models tend to improve recall, ranking, multi-behavior, and multi-domain performance on public data, while GPT and SASRec fail to scale, though the evidence lacks error bars.
-
FASTNav: Fine-tuned Adaptive Small-language-models Trained for Multi-point Robot Navigation
Fine-tuned small language models, coached by a GPT-4 teacher through iterative prompting, can perform multi-point robot navigation on edge devices with success rates approaching larger models.
-
Large Language models for Time Series Analysis: Techniques, Applications, and Challenges
A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.
Discussion (0). Continue with ORCID to comment.