Pith. sign in

Paper Citation Record · LEDGER

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models

As of 13 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.16991.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16991 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:45:20.262730Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact3
  • verified fuzzy18
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd5c9030-ce5b-43b1-a197-7b3d490c28c8 · outbound

This paper cites GPT-4 Technical Report.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.855397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.855397Z digest=sha256:f480199a469f1adc1d592763311697a22a5c8f28e98efe3d955f4b3180d8ce25

Observation 8c262949-5c80-4ffe-ab5d-dfb7786e1826 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.889753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.889753Z digest=sha256:07968caee6ba03683f37478d48bcdd42d5adcf0b8c2fa96e95038bf045c7a78d

Observation a0f75dc3-05f2-4fd2-9af7-18dbc227be31 · outbound

This paper cites RomeBERT: Robust Training of Multi-Exit BERT.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models RomeBERT: Robust Training of Multi-Exit BERT

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.900099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.900099Z digest=sha256:9fb2fb8beb4b53b4c01313ebdfa0dc2662f34d3e38920ae77c66cb653e95c102

Observation 15a08c71-bb43-4622-9580-fd57349bdcea · outbound

This paper cites Self-Knowledge Distillation in Natural Language Processing.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Self-Knowledge Distillation in Natural Language Processing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.910575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.910575Z digest=sha256:d572b8b3163d92e265d79a2b0f36f375f5de3b9bf17049ae97231ebfa6852140

Observation aae5d057-ae16-496e-8c74-ad02021637c9 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling the Knowledge in a Neural Network

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.927555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.927555Z digest=sha256:cf49017d3ffdece575c1be234fcfe7e4c9d47c8107c9bb5e591d235a6cb4a3c2

Observation 5ca3f85d-265e-4574-a822-8c505ac89bde · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.938293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.938293Z digest=sha256:2331cb267f3b6d1c01576baf785b2fbc99f1b91202049bc1a3f7fa388f55edb0

Observation e1d129ea-d400-4496-8c4c-eab8b38787fe · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.943841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.943841Z digest=sha256:a57421ef19e1638bd83935233ff3ace2f9cadc8511d29b450656c730562be186

Observation dbba784e-85c6-400e-8868-c4ba96e9c083 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models TinyBERT: Distilling BERT for Natural Language Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.949757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.949757Z digest=sha256:5800a4621da7601097802a580228d385745cfd7701d03649571fb142a6a63031

Observation ca37dec1-f373-4fd3-a31f-644000543747 · outbound

This paper cites DistiLLM: Towards Streamlined Distillation for Large Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DistiLLM: Towards Streamlined Distillation for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.955519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.955519Z digest=sha256:a982e7fcfbc7db7852c695a7b89d2ca301c2ef45c950232834247cf0a9b7402e

Observation 32436a6a-6055-4265-b736-69fcbe4a4017 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.960719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.960719Z digest=sha256:16f7d03e4a0b5319933245242a52ff2d865a5125a137508d6d578b47c33c4061

Observation 09e53d6a-af0e-4efb-9e38-eaa2008ab2fe · outbound

This paper cites A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:45:21.036127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:19.965765Z digest=sha256:b9719a55e775fbbde9c0ca63570e75eb342e5f947ea16a866fac331667c94b51

Observation 137651d1-b7d0-40f5-a61e-c143f1b4387d · outbound

This paper cites Dynamic Knowledge Distillation for Pre-trained Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Dynamic Knowledge Distillation for Pre-trained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.970608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.970608Z digest=sha256:2c893c8746a728357ae862253fdad5639c740c790bbbbe1fa02f4e590163381d

Observation bae96d79-ed68-40cc-b5d4-e0f678fbd458 · outbound

This paper cites Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Symbolic Chain-of-Thought Distillation: Small Models Can Also "Think" Step-by-Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.975270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.975270Z digest=sha256:191b3f1741013ee62b7a1a169e4314e6af620de73e99bcad34365f5f0181f6d9

Observation e5826e16-c35b-4d9e-b6f8-2fa2f3f3be66 · outbound

This paper cites HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained Transformers.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.980207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.980207Z digest=sha256:4ecd3b2374fa27625818c56d2304defd0f5ecd6a06e3b1d1e744ba6fcb256f89

Observation f263b0b1-ad39-454c-b65d-f4e05f8a9aed · outbound

This paper cites MixKD: Towards Efficient Distillation of Large-scale Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MixKD: Towards Efficient Distillation of Large-scale Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.986364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.986364Z digest=sha256:6a7653200bb5fc59cc8c68b12ef2b78eb8d9b1ea8a214ae7010c73aa7c943745

Observation 0664d6d8-4972-4801-8218-d7b2ae0f2091 · outbound

This paper cites A global past-future early exit method for accelerating inference of pre-trained language models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A global past-future early exit method for accelerating inference of pre-trained language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.745370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:19.991839Z digest=sha256:ec1e2f3a960c034102cb8d75e1efa87a68b0804c710a16a218827b1522aaf51f

Observation ed4711c8-ca6e-4016-b684-9f154eb58e4a · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.996512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.996512Z digest=sha256:6812c247495afd2ea3879822b8e8654621d13da07606ab5923dd81b099c26944

Observation 497eb5fe-b15e-4d14-8804-456ea0c82a86 · outbound

This paper cites FastBERT: a Self-distilling BERT with Adaptive Inference Time.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models FastBERT: a Self-distilling BERT with Adaptive Inference Time

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.001464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.001464Z digest=sha256:67d2cef6d696f823e2b7650519f55ff3f234b2d946c351f342defb8db600fd1b

Observation 7e8350f0-7a47-4d07-8541-a3d9be12c581 · outbound

This paper cites Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.007152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.007152Z digest=sha256:8b56c678b4d9b1ccbd4f1ae27c6a773633ff30c69bafd6021cf572d620fd1f66

Observation 5bbc6500-41bd-440b-87ee-34559a81b19b · outbound

This paper cites Multi-Task Deep Neural Networks for Natural Language Understanding.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Multi-Task Deep Neural Networks for Natural Language Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.012458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.012458Z digest=sha256:9ded85a1e05fc18f1e694320a0ca0b261ecb41bf9bb05c2f18437bf75c685183

Observation 9c22838a-2c52-4b74-91ac-eca9b5ce87f7 · outbound

This paper cites Big/little deep neural network for ultra low power inference.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Big/little deep neural network for ultra low power inference

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.728730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.018202Z digest=sha256:cf6cde1392d9d77f4269b3272b18e9b5aabf3a6aacd0f98ad2879f8e63717903

Observation ff5c6fb0-3b13-4d05-bc17-b544f3503d6f · outbound

This paper cites Distilling Linguistic Context for Language Model Compression.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Linguistic Context for Language Model Compression

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:45:20.844940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.023028Z digest=sha256:24d4d6af1efa30e576074b94591d623c07259923710fdd8ce9b0c8774cf3b988

Observation b017b3d7-60ae-48f4-bdcf-2939633300f4 · outbound

This paper cites Are NLP Models really able to Solve Simple Math Word Problems?.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Are NLP Models really able to Solve Simple Math Word Problems?

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.028487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.028487Z digest=sha256:d4f3b7d983330d087b42be302de0b4febd3ffc5dee4b750c4735e7db90e5beb6

Observation de8c9074-bee0-44a6-89f7-5569af0dcd49 · outbound

This paper cites WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.034338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.034338Z digest=sha256:f7a20c7757e828dcae6d7c3b8142a8f5a0db0051e6ceedde0489165ed9a0faee

Observation df5a3842-1519-497f-89c1-5d5e7005c112 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.039995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.039995Z digest=sha256:e3fcb71f4eb5d9af013cf7032b4ac3b6caad7e9b4079e467a76dbaf0895ec6f9

Observation 3b392747-60cc-4fb1-9045-90ebcced3b0e · outbound

This paper cites Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-12T12:45:20.753246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.045548Z digest=sha256:7411ec6d4e55fa1f31e1af1201cf128be95b68974da5eabdc39765735e2a734d

Observation c33a2cd8-3593-4078-b552-825e0089133d · outbound

This paper cites Choice of plausible alternatives: An evaluation of commonsense causal reasoning.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Choice of plausible alternatives: An evaluation of commonsense causal reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.052005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.052005Z digest=sha256:db8594098a8686a19fd8954bf744861e77ebfe27b7f20772ddc0c1fef765bed1

Observation ba3b25f5-f500-44ff-92aa-71bb5818687a · outbound

This paper cites Consistent Accelerated Inference via Confident Adaptive Transformers.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Consistent Accelerated Inference via Confident Adaptive Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.061438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.061438Z digest=sha256:1c6ac6de618513fdde35ba7c67aaf9ef946c6dce9ca870792d1f91a39025ad9a

Observation 9497d9c9-35e2-4518-8ebb-c7f59b4c34d0 · outbound

This paper cites The Right Tool for the Job: Matching Model and Instance Complexities.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models The Right Tool for the Job: Matching Model and Instance Complexities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.065896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.065896Z digest=sha256:f8d1ac27a47517969130a23786680ce3120235df5e9d45d320c18b3640748f31

Observation 62b38499-9f24-4bde-9d23-a12af74593a6 · outbound

This paper cites ResLoRA: Identity Residual Mapping in Low-Rank Adaption.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models ResLoRA: Identity Residual Mapping in Low-Rank Adaption

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.070951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.070951Z digest=sha256:137f23ac934420c271c7486217c9863e4e6a4d043e35b6cc23cdf7ffeaee944c

Observation cf79f997-8984-4103-93fc-7b75e0bf4c72 · outbound

This paper cites Distilling Reasoning Capabilities into Smaller Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Reasoning Capabilities into Smaller Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.076077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.076077Z digest=sha256:88637c53afcfea74bdef7c8ecfa5f0e65a660c528b656b1abd1b57c26d645922

Observation 671dc132-9b2f-4109-9e30-31bbe50d3824 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.080762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.080762Z digest=sha256:cceca309f81ea20fa69b72c777d31b733e0fb016aff09cdc5624f49e1238959f

Observation 83e6d587-5300-434b-9c34-677f6539318e · outbound

This paper cites Recursive deep models for semantic compositionality over a sentiment treebank.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recursive deep models for semantic compositionality over a sentiment treebank

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.085846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.085846Z digest=sha256:c509eaf425ec7fa9ce1e945e672749015f8837df295acd0bef85d2d0cfc236f5

Observation af6f62af-1040-45e5-baf7-f7e4e16d33ae · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Patient Knowledge Distillation for BERT Model Compression

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.090860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.090860Z digest=sha256:fabdc604ecbc2c2079c950ac4467f14a2f4f2eb72a0060d55f745c172f09918e

Observation 93a6a742-da9d-4b58-a6da-bed3ff595c0c · outbound

This paper cites Early Exiting with Ensemble Internal Classifiers.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Early Exiting with Ensemble Internal Classifiers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.095740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.095740Z digest=sha256:5e57529cd328b10f87442346679762fda3c897de287c9d84c1c3cb2d06f6d9fe

Observation 7ab00b3e-ea88-43c3-9cfb-b2f93fd69927 · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.100589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.100589Z digest=sha256:bf6f021ca94eee9d1764d3beab65f612f07756bbb9c5ba243104b58b773ddfa2

Observation 3304cc41-c9c9-425f-a353-6ce338a71649 · outbound

This paper cites Distilling Task-Specific Knowledge from BERT into Simple Neural Networks.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Distilling Task-Specific Knowledge from BERT into Simple Neural Networks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.105391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.105391Z digest=sha256:0f58b5799912a29e9c243d9cb3d2851599616f91f7b6bb2ea77e1478dbf32987

Observation ef10f0cc-f0bd-4634-9e45-833c2b964278 · outbound

This paper cites Branchynet: Fast inference via early exiting from deep neural networks.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Branchynet: Fast inference via early exiting from deep neural networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.110647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.110647Z digest=sha256:2ae6525bd5b03c5d4a5ce26d60a88f5dc075c6d7b9e0c1bf220d31633eec1f39

Observation bbf218b9-251a-4cee-b52a-19741c11c797 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.115991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.115991Z digest=sha256:5ac78d8b365e8d8b5123b70ab9ee54f6f08b06876ae87310c9fa9ff500e6cfe7

Observation 37cc3924-0318-48e3-ab6a-de92fcc986a5 · outbound

This paper cites Well-Read Students Learn Better: On the Importance of Pre-training Compact Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Well-Read Students Learn Better: On the Importance of Pre-training Compact Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.120970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.120970Z digest=sha256:06ba9a211d38e72983cfb9d3e4fab5c55ca7fc1a7b82a246d431c164142dde74

Observation d533196e-2c69-47ea-933c-03c5e95b23fd · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.125827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.125827Z digest=sha256:e30b142c290264383511a0c7efa26969f16b14eca48becbab90b31004ac705e6

Observation b95498d3-cdbf-4545-b8c0-ba3788c4c058 · outbound

This paper cites Improved Knowledge Distillation for Pre-trained Language Models via Knowledge Selection.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Improved Knowledge Distillation for Pre-trained Language Models via Knowledge Selection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.130774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.130774Z digest=sha256:80d360409889c89740e1c7338225f73d58fc5d2460eb4e5e4b705d0a5da6418f

Observation 8ae8f840-ff82-47e5-9918-91296ec94440 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.135954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.135954Z digest=sha256:2138678b0d832f58d2310287bc9dfb5f1074d1713fe51e8538841c1c5b82f8bd

Observation d18eb5ee-1628-485e-97f8-3d16e6c4981b · outbound

This paper cites Causal Distillation for Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Causal Distillation for Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.145677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.145677Z digest=sha256:50ca479598802de91081e2cefcdeec256038f2b726678e1f279f2d9c21f454ad

Observation 03018741-bede-4ba4-82b0-6921d519b596 · outbound

This paper cites Sparse Teachers Can Be Dense with Knowledge.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Sparse Teachers Can Be Dense with Knowledge

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.150812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.150812Z digest=sha256:24991ebe91e171a699a25c86c2d6291837aa4da4cd4454a23b904df8de32be66

Observation 1c1973a7-5fc7-4275-8667-be40f58c6934 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.155805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.155805Z digest=sha256:04eeaa63fb62935121042672a0158c8a1635c16f03b08ef269a9e190c50df9ed

Observation 38c8cdfb-2aef-490d-8854-d78295a9b47f · outbound

This paper cites Mini- mal distillation schedule for extreme language model compression.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Mini- mal distillation schedule for extreme language model compression

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.682340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.160870Z digest=sha256:cd5fc4bb31ac9b00c5909d31f4d38c75cb9634b50933044ccdf1940242954bde

Observation 5443ab3d-21fd-4b53-98f2-b7b7a4a15b2a · outbound

This paper cites Small Language Models Need Strong Verifiers to Self-Correct Reasoning.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Small Language Models Need Strong Verifiers to Self-Correct Reasoning

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.165403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.165403Z digest=sha256:1ded1747e6d9aa5f87410becc81e961686e74d111b3db05f07f74d3c984d15e7

Observation e95a2405-d66b-47ba-83af-749879a5d814 · outbound

This paper cites A Survey of Large Language Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models A Survey of Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.170533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.170533Z digest=sha256:56103e650cb6c4e7b1b36f4d59288ab59827162d9aa4df38d3960fd0e4faae43

Observation e0daff95-ede8-4958-ae03-03b63f0ef06e · outbound

This paper cites BERT Learns to Teach: Knowledge Distillation with Meta Learning.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BERT Learns to Teach: Knowledge Distillation with Meta Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.175736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.175736Z digest=sha256:300e3f3ca024841a48f2ac5eae079ec81f3a02ea8a335df4871d84edaa0f206c

Observation 6597f0b2-09ce-435a-8e16-6b832946d65a · outbound

This paper cites PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models PaD: Program-aided Distillation Can Teach Small Models Reasoning Better than Chain-of-thought Fine-tuning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.180921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.180921Z digest=sha256:4a573fea26ad6e52d219c6a3222e3f677ccb14c17625040ce832946ac2fb79bb

Observation 9cb60f6c-2414-499c-be8b-fce08240fe30 · outbound

This paper cites Recently, researchers have attached great significance to the KD study in PLMs (Sun et al., 2022).

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recently, researchers have attached great significance to the KD study in PLMs (Sun et al., 2022)

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.666937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.186087Z digest=sha256:a00ed3ef0c311901aa4ec79e6af2cdf62fdde843c8f2fe213afd1ac79cde0ae1

Observation 6d36fa64-e71f-4c50-8a27-d7d38b4a00aa · outbound

This paper cites an unresolved cited work.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:45:21.652060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.190818Z digest=sha256:7a9978e28af9c67f715b67fb060b549b22a3d9fb3712d5debb5ffd671415514b

Observation b85c0e20-680d-47d1-b587-3cf7a841d29d · outbound

This paper cites Recently, inspired by MetaDistil (Zhou et al., 2021), Ren et al.(Ren et al.,.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Recently, inspired by MetaDistil (Zhou et al., 2021), Ren et al.(Ren et al.,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.637067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.195972Z digest=sha256:b0a8a84ea343d6dc9824dc692da5c123039041373716730b0a3b9fb18aadbef2

Observation 8b50b7c1-5779-43ae-8b14-7b1cd17ebb66 · outbound

This paper cites For instance, DistilBERT (Sanh et al., 2019), MINILM (Wang et al., 2020b), and MobileBERT (Sun et al.,.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models For instance, DistilBERT (Sanh et al., 2019), MINILM (Wang et al., 2020b), and MobileBERT (Sun et al.,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.621305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.200598Z digest=sha256:d4ab130020b0694f3b96f49320a12a4d7c6f59c47fbdacc943258a18bb791c74

Observation 21641f7e-1506-47f9-93f0-d3a635b199d4 · outbound

This paper cites Based on MiniLLM, works on studying KD for auto-regressive LLMs (Agarwal et al., 2024; Ko et al.,.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Based on MiniLLM, works on studying KD for auto-regressive LLMs (Agarwal et al., 2024; Ko et al.,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.605112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.205290Z digest=sha256:db605f33e6e01f33a33c39c05585a056d5c001939aabb24291820d3ac47869ef

Observation 3b573a0e-be15-4dd0-bb8c-7f8858b78e7e · outbound

This paper cites an unresolved cited work.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:45:21.589209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.209904Z digest=sha256:6af7c7674b33546f28bc5b50e371eb5ab6b1aba158dfca82576ab4fd095e76fd

Observation 9ab00848-613e-4c2d-8256-5239697c4eef · outbound

This paper cites With the emergence of PLMs (Sun et al., 2022), people have begun to study the application of SelfD on them.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models With the emergence of PLMs (Sun et al., 2022), people have begun to study the application of SelfD on them

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.574150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.214947Z digest=sha256:272872c4c0c2b4eade9f1f8df78410aefa1eda2411cfae89c1bcaceb9ae9f16c

Observation 5d50122c-5600-4942-a0cf-eb39fedb3bfe · outbound

This paper cites Though not reducing model sizes, it decreases computation by using inserted internal classifiers into a Transformer-based model (e.g., 12-layer BERT-base).

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Though not reducing model sizes, it decreases computation by using inserted internal classifiers into a Transformer-based model (e.g., 12-layer BERT-base)

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.558362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.219614Z digest=sha256:d9d49560519f73140b0d47c5be712a7f68667d050b2a602afba596aa9776f063

Observation 7d47cb61-60fd-4e7a-994c-ab4c20bc8ad2 · outbound

This paper cites EE techniques for PLMs focus on exit criteria, which currently have three types (Xu & McAuley, 2023): confidence estimation, internal ensemble, and learning to exit.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models EE techniques for PLMs focus on exit criteria, which currently have three types (Xu & McAuley, 2023): confidence estimation, internal ensemble, and learning to exit

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.542281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.224375Z digest=sha256:52bc834fc25c1b38d441e6d12ec644d47c0ac9cf6e5a257a2702ac3c7b22dffd

Observation 7c13ef59-3c99-43fb-94b4-03bb515b71ab · outbound

This paper cites During inference, the model exits early when an IC predicts a probability with an entropy below the threshold.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models During inference, the model exits early when an IC predicts a probability with an entropy below the threshold

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.525746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.229054Z digest=sha256:3fd77cdb507bee0b0107bd379369ed0b7218ab8db40fd45ea69e2ad5775ff3e1

Observation f2f74490-4a1e-4196-a3f0-22b74700243c · outbound

This paper cites an unresolved cited work.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:45:21.505162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.233544Z digest=sha256:509e6a71d424079af283d3d5d0374238d6c3b5b9b92411447b93d2b7464f5c11

Observation f4ad7e23-6c2f-4141-a278-bcadebe28b09 · outbound

This paper cites Liao et al.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Liao et al

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.486973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.238204Z digest=sha256:e7287d657007bae42dc0ff194d01d1ebb1a7c812d62f3b6577f685ee5855be69

Observation ede566cf-a38f-4e50-8f05-ec132f8c86b0 · outbound

This paper cites It serves as a benchmark for evaluating the performance of models across various language understanding tasks.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It serves as a benchmark for evaluating the performance of models across various language understanding tasks

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.466806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.242977Z digest=sha256:6636fe8783067525c2d28816a1ae07622c50b22e0296789c3b755a2df068dfb0

Observation 0b4d1b0e-e4d4-46eb-9f52-eed6ce9b2ee7 · outbound

This paper cites It was introduced as a more challenging successor to the original GLUE benchmark (Wang et al., 2018), reflecting the rapid advancements in NLP technologies and model capabilities.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It was introduced as a more challenging successor to the original GLUE benchmark (Wang et al., 2018), reflecting the rapid advancements in NLP technologies and model capabilities

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.451110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.247909Z digest=sha256:23c4f1db89da6b03c2723ecdd7f9c15f22de23c542c8dc6e6d3ab1e6b95f25a0

Observation d2f3c1f6-c541-4004-9e86-74765380d621 · outbound

This paper cites For commonsense tasks, we select HellaSwag (HS) (Zellers et al., 2019).

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models For commonsense tasks, we select HellaSwag (HS) (Zellers et al., 2019)

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.435163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.252812Z digest=sha256:a44dd4336ab734cfcf75eea1cfc1cabf668d21167d2da613c26834c1db9bc6ec

Observation e84196c9-eda2-401c-9269-f86c2006c694 · outbound

This paper cites This dataset is designed to test both comprehension and arithmetic skills in a more controlled synthetic setting.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models This dataset is designed to test both comprehension and arithmetic skills in a more controlled synthetic setting

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.418668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.257624Z digest=sha256:cf6b31ce6c70ce22886772bd5d345275d21b844c89b83774d7a99ba869d528cc

Observation 524d002e-6e82-423f-8ad2-46ef66017154 · outbound

This paper cites It presents contexts from a wide array of domains and requires models to predict the most likely or plausible continuation among given choices.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models It presents contexts from a wide array of domains and requires models to predict the most likely or plausible continuation among given choices

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.401539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:20.262730Z digest=sha256:10d166945d69eed47d1ff21361f7853d82d62a095a62fc74e87771a0e583b8c7

Observation 23666cba-034c-4974-b035-d8916a2e660e · outbound

This paper cites MCC-KD: Multi-CoT Consistent Knowledge Distillation.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MCC-KD: Multi-CoT Consistent Knowledge Distillation

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.866920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.866920Z digest=sha256:585d807cefcde952d309e20b05e7132a170276668d6b21a5b4f3dd0beed31d16

Observation adc8b3f8-4e80-46d2-894b-d60ba539272b · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.056612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.056612Z digest=sha256:d3ee770afa6674c0eea6e74be679e51ab5cf329ad4699dc6c76dfed87142e61d

Observation 1353c1cf-bae0-4b15-b3c8-b866e30d01ce · outbound

This paper cites Large Language Models Are Reasoning Teachers.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Large Language Models Are Reasoning Teachers

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.932907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.932907Z digest=sha256:35c85ddbfa36fb6293a05c380e1ab3689f775e40fc9a86666b5ab21e05c038f7

Observation 09fe76da-f4e4-4059-8432-c263093fe682 · outbound

This paper cites DeBERTa: Decoding-enhanced BERT with Disentangled Attention.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.916845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.916845Z digest=sha256:f18825fea1a40b5d41b2b144a254c6d6b14387fd3ba1aa8511926bd2aa310bd7

Observation 22f5248d-ac3c-449a-8b1f-7997f5b05802 · outbound

This paper cites One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:20.140755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:20.140755Z digest=sha256:fd326dcbea3aafd7e35965a46cdd55164224631331d086b80de71121a4bb6951

Observation 05589fc6-7fef-492d-ac64-59311a9f490c · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.872344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.872344Z digest=sha256:8e81ac3d7220684672a6a92af72971d545dc7ba300ee59ef4a70b5da5e8c4412

Observation e4f803f6-a2d3-47c7-a113-2e37f89f7a62 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Training Verifiers to Solve Math Word Problems

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.878101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.878101Z digest=sha256:b0e9c6fb691f34b7bafacf84b47606598ae3954f51c88f0efd1beeaa16065c3b

Observation 13a62572-9a71-48ec-83e5-754146d4fd36 · outbound

This paper cites DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models DeBERTaV3: Improving DeBERTa using ELECTRA-Style Pre-Training with Gradient-Disentangled Embedding Sharing

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.922552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.922552Z digest=sha256:dccbf556a5370e6039f5d4ff085e78d8997a81e531f99a7b4ba7ab23aa5c0791

Observation 5352a212-24e8-4e85-9c1f-465713e87e30 · outbound

This paper cites Cost-effective distillation of large language models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Cost-effective distillation of large language models

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:45:21.761911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T12:45:19.884636Z digest=sha256:ebae27b61c9bd7b47acd4d6814be34d5fa1eb4c9236cae3426a4cf9f55474d4c

Observation cf9ea5c9-77d5-49c9-bc3e-6bad42e70709 · outbound

This paper cites The Llama 3 Herd of Models.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models The Llama 3 Herd of Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.894833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.894833Z digest=sha256:970271bfb9c64e223f70b1f3a0b7927ca801659dae6098a2a800e2077f745674

Observation 45ca44a1-94d7-4a30-a9dd-e1a56e2ee52b · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models Reinforced Self-Training (ReST) for Language Modeling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.905572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.905572Z digest=sha256:ddacc4e068c95afbc4506dd21c40a4e422cb654f946d1a512e42f34fa18268e3

Observation bc63d689-db93-4c23-96e8-aedbe1e2b784 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T12:45:19.861361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:45:19.861361Z digest=sha256:da7167551b53794cba47957b23e4020705f770c630000db37b3927eecc1f4242

Pith citing papers

No inbound Pith citation observations are available.