Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:45:20.262730Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.16991.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:45:20.262730Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
80 of 80 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd5c9030-ce5b-43b1-a197-7b3d490c28c8 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c262949-5c80-4ffe-ab5d-dfb7786e1826 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0f75dc3-05f2-4fd2-9af7-18dbc227be31 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models RomeBERT: Robust Training of Multi-Exit BERT
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a08c71-bb43-4622-9580-fd57349bdcea · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Self-Knowledge Distillation in Natural Language Processing
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae5d057-ae16-496e-8c74-ad02021637c9 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling the Knowledge in a Neural Network
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca3f85d-265e-4574-a822-8c505ac89bde · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1d129ea-d400-4496-8c4c-eab8b38787fe · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbba784e-85c6-400e-8868-c4ba96e9c083 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models TinyBERT: Distilling BERT for Natural Language Understanding
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca37dec1-f373-4fd3-a31f-644000543747 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32436a6a-6055-4265-b736-69fcbe4a4017 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09e53d6a-af0e-4efb-9e38-eaa2008ab2fe · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 137651d1-b7d0-40f5-a61e-c143f1b4387d · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Dynamic Knowledge Distillation for Pre-trained Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bae96d79-ed68-40cc-b5d4-e0f678fbd458 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5826e16-c35b-4d9e-b6f8-2fa2f3f3be66 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f263b0b1-ad39-454c-b65d-f4e05f8a9aed · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MixKD: Towards Efficient Distillation of Large-scale Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0664d6d8-4972-4801-8218-d7b2ae0f2091 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A global past-future early exit method for accelerating inference of pre-trained language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed4711c8-ca6e-4016-b684-9f154eb58e4a · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497eb5fe-b15e-4d14-8804-456ea0c82a86 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models FastBERT: a Self-distilling BERT with Adaptive Inference Time
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e8350f0-7a47-4d07-8541-a3d9be12c581 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bbc6500-41bd-440b-87ee-34559a81b19b · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Multi-Task Deep Neural Networks for Natural Language Understanding
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c22838a-2c52-4b74-91ac-eca9b5ce87f7 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Big/little deep neural network for ultra low power inference
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff5c6fb0-3b13-4d05-bc17-b544f3503d6f · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Linguistic Context for Language Model Compression
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b017b3d7-60ae-48f4-bdcf-2939633300f4 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Are NLP Models really able to Solve Simple Math Word Problems?
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8c9074-bee0-44a6-89f7-5569af0dcd49 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5a3842-1519-497f-89c1-5d5e7005c112 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models SQuAD: 100,000+ Questions for Machine Comprehension of Text
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b392747-60cc-4fb1-9045-90ebcced3b0e · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c33a2cd8-3593-4078-b552-825e0089133d · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3b25f5-f500-44ff-92aa-71bb5818687a · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Consistent Accelerated Inference via Confident Adaptive Transformers
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9497d9c9-35e2-4518-8ebb-c7f59b4c34d0 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models The Right Tool for the Job: Matching Model and Instance Complexities
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62b38499-9f24-4bde-9d23-a12af74593a6 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models ResLoRA: Identity Residual Mapping in Low-Rank Adaption
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf79f997-8984-4103-93fc-7b75e0bf4c72 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Reasoning Capabilities into Smaller Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 671dc132-9b2f-4109-9e30-31bbe50d3824 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83e6d587-5300-434b-9c34-677f6539318e · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recursive deep models for semantic compositionality over a sentiment treebank
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af6f62af-1040-45e5-baf7-f7e4e16d33ae · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Patient Knowledge Distillation for BERT Model Compression
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a6a742-da9d-4b58-a6da-bed3ff595c0c · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Early Exiting with Ensemble Internal Classifiers
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab00b3e-ea88-43c3-9cfb-b2f93fd69927 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3304cc41-c9c9-425f-a353-6ce338a71649 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef10f0cc-f0bd-4634-9e45-833c2b964278 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Branchynet: Fast inference via early exiting from deep neural networks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbf218b9-251a-4cee-b52a-19741c11c797 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37cc3924-0318-48e3-ab6a-de92fcc986a5 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d533196e-2c69-47ea-933c-03c5e95b23fd · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b95498d3-cdbf-4545-b8c0-ba3788c4c058 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Improved Knowledge Distillation for Pre-trained Language Models via Knowledge Selection
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae8f840-ff82-47e5-9918-91296ec94440 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d18eb5ee-1628-485e-97f8-3d16e6c4981b · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Causal Distillation for Language Models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03018741-bede-4ba4-82b0-6921d519b596 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Sparse Teachers Can Be Dense with Knowledge
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c1973a7-5fc7-4275-8667-be40f58c6934 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38c8cdfb-2aef-490d-8854-d78295a9b47f · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Mini- mal distillation schedule for extreme language model compression
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5443ab3d-21fd-4b53-98f2-b7b7a4a15b2a · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Small Language Models Need Strong Verifiers to Self-Correct Reasoning
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95a2405-d66b-47ba-83af-749879a5d814 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Survey of Large Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0daff95-ede8-4958-ae03-03b63f0ef06e · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BERT Learns to Teach: Knowledge Distillation with Meta Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6597f0b2-09ce-435a-8e16-6b832946d65a · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb60f6c-2414-499c-be8b-fce08240fe30 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recently, researchers have attached great significance to the KD study in PLMs (Sun et al., 2022)
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d36fa64-e71f-4c50-8a27-d7d38b4a00aa · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b85c0e20-680d-47d1-b587-3cf7a841d29d · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recently, inspired by MetaDistil (Zhou et al., 2021), Ren et al.(Ren et al.,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b50b7c1-5779-43ae-8b14-7b1cd17ebb66 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models For instance, DistilBERT (Sanh et al., 2019), MINILM (Wang et al., 2020b), and MobileBERT (Sun et al.,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 21641f7e-1506-47f9-93f0-d3a635b199d4 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Based on MiniLLM, works on studying KD for auto-regressive LLMs (Agarwal et al., 2024; Ko et al.,
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b573a0e-be15-4dd0-bb8c-7f8858b78e7e · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ab00848-613e-4c2d-8256-5239697c4eef · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models With the emergence of PLMs (Sun et al., 2022), people have begun to study the application of SelfD on them
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d50122c-5600-4942-a0cf-eb39fedb3bfe · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Though not reducing model sizes, it decreases computation by using inserted internal classifiers into a Transformer-based model (e.g., 12-layer BERT-base)
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7d47cb61-60fd-4e7a-994c-ab4c20bc8ad2 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models EE techniques for PLMs focus on exit criteria, which currently have three types (Xu & McAuley, 2023): confidence estimation, internal ensemble, and learning to exit
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c13ef59-3c99-43fb-94b4-03bb515b71ab · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models During inference, the model exits early when an IC predicts a probability with an entropy below the threshold
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f2f74490-4a1e-4196-a3f0-22b74700243c · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4ad7e23-6c2f-4141-a278-bcadebe28b09 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Liao et al
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ede566cf-a38f-4e50-8f05-ec132f8c86b0 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It serves as a benchmark for evaluating the performance of models across various language understanding tasks
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b4d1b0e-e4d4-46eb-9f52-eed6ce9b2ee7 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It was introduced as a more challenging successor to the original GLUE benchmark (Wang et al., 2018), reflecting the rapid advancements in NLP technologies and model capabilities
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2f3c1f6-c541-4004-9e86-74765380d621 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models For commonsense tasks, we select HellaSwag (HS) (Zellers et al., 2019)
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e84196c9-eda2-401c-9269-f86c2006c694 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models This dataset is designed to test both comprehension and arithmetic skills in a more controlled synthetic setting
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 524d002e-6e82-423f-8ad2-46ef66017154 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It presents contexts from a wide array of domains and requires models to predict the most likely or plausible continuation among given choices
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23666cba-034c-4974-b035-d8916a2e660e · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MCC-KD: Multi-CoT Consistent Knowledge Distillation
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adc8b3f8-4e80-46d2-894b-d60ba539272b · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1353c1cf-bae0-4b15-b3c8-b866e30d01ce · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Large Language Models Are Reasoning Teachers
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09fe76da-f4e4-4059-8432-c263093fe682 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DeBERTa: Decoding-enhanced BERT with Disentangled Attention
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22f5248d-ac3c-449a-8b1f-7997f5b05802 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05589fc6-7fef-492d-ac64-59311a9f490c · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f803f6-a2d3-47c7-a113-2e37f89f7a62 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Training Verifiers to Solve Math Word Problems
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a62572-9a71-48ec-83e5-754146d4fd36 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5352a212-24e8-4e85-9c1f-465713e87e30 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Cost-effective distillation of large language models
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf9ea5c9-77d5-49c9-bc3e-6bad42e70709 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models The Llama 3 Herd of Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ca44a1-94d7-4a30-a9dd-e1a56e2ee52b · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Reinforced Self-Training (ReST) for Language Modeling
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc63d689-db93-4c23-96e8-aedbe1e2b784 · outbound
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.