Pith. sign in

Paper Citation Record · LEDGER

Efficient Knowledge Injection in LLMs via Self-Distillation

As of 12 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 7 inbound Pith citation observations for arXiv:2412.14964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14964 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:52:37.270748Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:04:31.647548Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:56:47.730179Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03e56360-c8f8-4595-8e27-5336625d58f7 · outbound

This paper cites GPT-4 Technical Report.

Efficient Knowledge Injection in LLMs via Self-Distillation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.086452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.086452Z digest=sha256:ade1a057f070422a00f520c8ab9dbb914836e9cd9b475c066a3605e12f9d94b3

Observation 47d91a7d-ffe6-4e50-8395-36040de64166 · outbound

This paper cites Adapting Language Models to Compress Contexts.

Efficient Knowledge Injection in LLMs via Self-Distillation Adapting Language Models to Compress Contexts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.104311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.104311Z digest=sha256:0903b626302ac6b2401ae3d022caf5bf2c6104c15077bc3365e7f05d099b3e9d

Observation 56089bf7-56f8-46e6-a1b0-93035819f9a5 · outbound

This paper cites The Llama 3 Herd of Models.

Efficient Knowledge Injection in LLMs via Self-Distillation The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.112713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.112713Z digest=sha256:99a86e7083aaf30a13a333cbc9cb2936910ae6d01e507bb2371d4eff93ac949a

Observation 7470d0fb-9905-4fb6-9e1e-93189fe3665b · outbound

This paper cites The False Promise of Imitating Proprietary LLMs.

Efficient Knowledge Injection in LLMs via Self-Distillation The False Promise of Imitating Proprietary LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.116782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.116782Z digest=sha256:9db9594295ec93fb6a4e2ef8b64c2b8cc8877095e4316a3cfd0a4bd3ff6f5413

Observation 7796b37f-3803-4045-92a5-8329e84c5a90 · outbound

This paper cites Don't Stop Pretraining: Adapt Language Models to Domains and Tasks.

Efficient Knowledge Injection in LLMs via Self-Distillation Don't Stop Pretraining: Adapt Language Models to Domains and Tasks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.125237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.125237Z digest=sha256:4c982f360a503ca304811d168dc1bd7573e8072fafee6ea2776202f91fb18796

Observation 9578ef48-9612-46c7-8956-972139b897bf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Efficient Knowledge Injection in LLMs via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.134081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.134081Z digest=sha256:d69a02fd04bf8247c5f1af8a7446f6a2eadc374cc49776c6fc2ccd1ab0fbfcfc

Observation 6126f122-a95f-4b8c-982d-f151609a3ea6 · outbound

This paper cites RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation RAVEN: In-Context Learning with Retrieval-Augmented Encoder-Decoder Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.143760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.143760Z digest=sha256:a9af058db0471ad36ce5417bac3f0eda22f05c5a41a7e1d595d1b415ab87e791

Observation 2d269c3c-e99d-432c-9a0b-87bbf3a0aaa8 · outbound

This paper cites Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering.

Efficient Knowledge Injection in LLMs via Self-Distillation Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.148846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.148846Z digest=sha256:fb22aff2a3971766e6fe603b933915ffc8bafb82ef30a9c23272d087c8cf4c19

Observation 805153a0-8015-4b22-91e7-c5dee28c07b2 · outbound

This paper cites Generalization through Memorization: Nearest Neighbor Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Generalization through Memorization: Nearest Neighbor Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.153376Z digest=sha256:c534fbc013e92511b9fdcfc4e4c2376b738e31ff4b4d515b86828f81924d51df

Observation 763a6aab-fa9c-41fb-ac48-d46c6e206bb6 · outbound

This paper cites RA-DIT: Retrieval-Augmented Dual Instruction Tuning.

Efficient Knowledge Injection in LLMs via Self-Distillation RA-DIT: Retrieval-Augmented Dual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.162226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.162226Z digest=sha256:ff80d33d0cf57d7525c538034d9152df88127351aea02a992e2e76afd5b7af17

Observation e457e219-f634-44d5-a729-f0ae9a5cce2f · outbound

This paper cites Structure-aware Domain Knowledge Injection for Large Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Structure-aware Domain Knowledge Injection for Large Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:52:37.522227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.167306Z digest=sha256:5a4b03aecc409e6c3ffc38488c0164776cb9f2352a4e01115db12dd65800a640

Observation ef842316-cf51-4bb9-a7ce-5973f06ce800 · outbound

This paper cites ChatQA: Surpassing GPT-4 on Conversational QA and RAG.

Efficient Knowledge Injection in LLMs via Self-Distillation ChatQA: Surpassing GPT-4 on Conversational QA and RAG

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.171818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.171818Z digest=sha256:2529b8e1165ff5ec18eddac36396310eaa61f7814790a0b85bb09fa89c2a8c44

Observation d9b8178e-5f41-4a8a-ab2b-ab4c75d8a82d · outbound

This paper cites When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories.

Efficient Knowledge Injection in LLMs via Self-Distillation When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.176676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.176676Z digest=sha256:46e7d37591d15d12fef21f962c095b10238adb673492e574122c2f71698c948d

Observation 03a92b77-3952-4e84-b964-e741a3d76cf3 · outbound

This paper cites Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning.

Efficient Knowledge Injection in LLMs via Self-Distillation Injecting New Knowledge into Large Language Models via Supervised Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.181478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.181478Z digest=sha256:66fd7ba09be7d3db52c86bd5483e2850fad42630266be57007381a69a6ab2162

Observation e55a76c8-fa4c-4235-b91f-3c9b16d18ec2 · outbound

This paper cites Orca 2: Teaching Small Language Models How to Reason.

Efficient Knowledge Injection in LLMs via Self-Distillation Orca 2: Teaching Small Language Models How to Reason

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.186069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.186069Z digest=sha256:5a7076035b508f033e1ed1322f994dd99a5dd1665a43f8b338a0e3b5247ff58f

Observation 26f47e3f-fd91-49cd-95c9-4cc7954909a3 · outbound

This paper cites XtremeDistil: Multi-stage Distillation for Massive Multilingual Models.

Efficient Knowledge Injection in LLMs via Self-Distillation XtremeDistil: Multi-stage Distillation for Massive Multilingual Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.191047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.191047Z digest=sha256:b1f7aab68e30e6e5b3ca86195e66045f161e1b7e8867fc7efd49001b773670db

Observation 7fe94a89-d997-4c74-b92f-321de45756cd · outbound

This paper cites Orca: Progressive Learning from Complex Explanation Traces of GPT-4.

Efficient Knowledge Injection in LLMs via Self-Distillation Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.195772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.195772Z digest=sha256:1934d369c2999090b63377fb87545add1e1a2b46885da63dfff43eaf058c76dd

Observation a180c0bd-0333-46ef-aed6-f19a61b16814 · outbound

This paper cites Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation.

Efficient Knowledge Injection in LLMs via Self-Distillation Learning to Generate Instruction Tuning Datasets for Zero-Shot Task Adaptation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.200189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.200189Z digest=sha256:3b3ccb6dfdef34f15e56e0911e8371fb861175ead7a24f166294d5f5a11000cf

Observation 46ae1da7-bb1f-4efc-9158-6a63d3e710b9 · outbound

This paper cites Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs.

Efficient Knowledge Injection in LLMs via Self-Distillation Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.204758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.204758Z digest=sha256:9ba83b0904205818f541cf454edd058b3b06f18ba756c605c0549e8902bf4032

Observation a7e5019a-b4ba-410f-b011-5f50298570a3 · outbound

This paper cites Instruction Tuning with GPT-4.

Efficient Knowledge Injection in LLMs via Self-Distillation Instruction Tuning with GPT-4

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.209209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.209209Z digest=sha256:44d14f1e497308823f2bba88c76cf457ad83c9ee2a2ca70663ab46e4191cb4f9

Observation 1154c72e-e39c-4a94-93eb-bcfe8809d26c · outbound

This paper cites In-Context Editing: Learning Knowledge from Self-Induced Distributions.

Efficient Knowledge Injection in LLMs via Self-Distillation In-Context Editing: Learning Knowledge from Self-Induced Distributions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.213565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.213565Z digest=sha256:29a39c093f02831ea38532e0f8a4623ba1143760ac83247b5928382699152480

Observation 9f8816d2-bda9-4d46-8d19-94c6926739ec · outbound

This paper cites Multitask Prompted Training Enables Zero-Shot Task Generalization.

Efficient Knowledge Injection in LLMs via Self-Distillation Multitask Prompted Training Enables Zero-Shot Task Generalization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.222089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.222089Z digest=sha256:24b79265002b401b44c0d9b43ba0e23b965c6d24d9f8132183a8df237ece11c0

Observation a716a0cc-f600-487b-afc5-c2031fd2c167 · outbound

This paper cites REPLUG: Retrieval-Augmented Black-Box Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation REPLUG: Retrieval-Augmented Black-Box Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.226208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.226208Z digest=sha256:97c734b4d7f4a5a3e994c1777e05430e797d28453d84686e824a17a7d3f210f9

Observation ed94ebae-474a-4acb-a0d6-1e8c7e463793 · outbound

This paper cites InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining.

Efficient Knowledge Injection in LLMs via Self-Distillation InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.230237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.230237Z digest=sha256:210b4a42f0c51f848fa978a230f49bfdfec843ee8a0b71a97a6470a1a85eb894

Observation 6a66c2b8-fb06-43a0-8531-eda96ea69d6c · outbound

This paper cites In-Context Former: Lightning-fast Compressing Context for Large Language Model.

Efficient Knowledge Injection in LLMs via Self-Distillation In-Context Former: Lightning-fast Compressing Context for Large Language Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.234248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.234248Z digest=sha256:4edc351a6ab9918921523e9c2a4ded3745873d42dbb3cfd6ab1bc60c7e41e855

Observation 923c68cc-4294-4dd4-a31e-c97c1eb74b92 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Efficient Knowledge Injection in LLMs via Self-Distillation HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.238593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.238593Z digest=sha256:a4d92be0341a186e585f71bd9b697187df76f6372c6ebf67c78906fd2b81cc21

Observation dbb9e362-91b0-48c8-ab4d-527e837e71b0 · outbound

This paper cites RAFT: Adapting Language Model to Domain Specific RAG.

Efficient Knowledge Injection in LLMs via Self-Distillation RAFT: Adapting Language Model to Domain Specific RAG

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.246602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.246602Z digest=sha256:77967b4d75861bbe10b69892c13b907b96df830cfd76013d9465859396060472

Observation 41bedcf0-75ff-4798-ae30-f499cb4282e0 · outbound

This paper cites an unresolved cited work.

Efficient Knowledge Injection in LLMs via Self-Distillation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:52:37.826001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.250250Z digest=sha256:622246574adfad0bf0bd294de877546ad3fac5a3cb81228953392d497a947bd6

Observation 04e9f70e-b46c-466b-8655-729caef3457b · outbound

This paper cites C Related Work: More Detailed Review of Context Distillation In prior work, context distillation has been used for in-context learning and qualitatively modifying LLM behavior.

Efficient Knowledge Injection in LLMs via Self-Distillation C Related Work: More Detailed Review of Context Distillation In prior work, context distillation has been used for in-context learning and qualitatively modifying LLM behavior

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.811775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.254877Z digest=sha256:6a03a9af5028bae09db795c1d468e1b412ebeea60b8cc7974f0bb5c0bd17e347

Observation 9cbae46e-8cd4-4a99-91cb-f11120cdc24f · outbound

This paper cites We found Bonito capable of generating competitive questions for the New York Times dataset.

Efficient Knowledge Injection in LLMs via Self-Distillation We found Bonito capable of generating competitive questions for the New York Times dataset

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.796868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.258972Z digest=sha256:cff2cbce5bdf9e7cf997514c32fb18b3271d323fd9ee66c8073eb2921f44ac90

Observation 9b9bc418-12e8-4212-bca3-54286de8d69b · outbound

This paper cites To explore potential factors underlying this phenomenon, we examine two key statistical properties of the teacher model’s outputs:.

Efficient Knowledge Injection in LLMs via Self-Distillation To explore potential factors underlying this phenomenon, we examine two key statistical properties of the teacher model’s outputs:

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.783268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.262850Z digest=sha256:1258f8ae54533a0d5588195cc126edbdc4123c969499fbcf1dbb9e82a941b69a

Observation 84748132-158c-4def-b593-fbe1bb9e97ad · outbound

This paper cites This characteristic may help explain why Llama-3- 8B-Instruct demonstrates superior performance as an expert compared to Qwen2.5-72B-Instruct (Table 3).

Efficient Knowledge Injection in LLMs via Self-Distillation This characteristic may help explain why Llama-3- 8B-Instruct demonstrates superior performance as an expert compared to Qwen2.5-72B-Instruct (Table 3)

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T11:52:37.766198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.266686Z digest=sha256:33946ac571651f77b588ad343aac1e637e73f4c3aff9150038452f4898b0cdbe

Observation 241a108e-0ec2-4009-a715-e53238617d53 · outbound

This paper cites Splendid Cities.

Efficient Knowledge Injection in LLMs via Self-Distillation Splendid Cities

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:52:37.752900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:52:37.270748Z digest=sha256:be3928517257bc91506b75d8764450242b4a4eb8fa8e9bf37261e24cbf55c2d9

Observation 123ed824-d287-4346-b5ab-1066c915c06f · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Efficient Knowledge Injection in LLMs via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.217777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.217777Z digest=sha256:9adfe78ed797778e8eab9988f609c8be081526680fd78fed3f1bf7adcd34f5a5

Observation f5475e12-fc41-4d5d-a90e-3e6d776adc2f · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation LoRA: Low-Rank Adaptation of Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.138767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.138767Z digest=sha256:48392f86eb0a841c377bd825b60d405531d28af1237baf81713f9c81380c9d7e

Observation 8b57b03b-a881-4fed-b327-2bbbb52cb362 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Efficient Knowledge Injection in LLMs via Self-Distillation Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.157917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.157917Z digest=sha256:e2aa68fc175d18464528b38921275e5269ab43e001b91cdc1f08f72ddffa5504

Observation 54ab17d3-19ec-4103-9bdb-3bbfe2983a89 · outbound

This paper cites Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model.

Efficient Knowledge Injection in LLMs via Self-Distillation Qilin-Med: Multi-stage Knowledge Injection Advanced Medical Large Language Model

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.242707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.242707Z digest=sha256:0c09c606d4415ff5315c092ab6a300aa0dc5b6f814f643c0542955575e94c274

Observation 072dd78b-6d67-4811-a065-c09d41319785 · outbound

This paper cites Prompt Injection: Parameterization of Fixed Inputs.

Efficient Knowledge Injection in LLMs via Self-Distillation Prompt Injection: Parameterization of Fixed Inputs

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.108513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.108513Z digest=sha256:72b6d653b4eefb6ff1c716a6fdf570ff1e2b5486fe361049ec9688634c065695

Observation 992cede1-a7a6-4848-b9c0-527f3f93f629 · outbound

This paper cites Revisiting Self-Training for Neural Sequence Generation.

Efficient Knowledge Injection in LLMs via Self-Distillation Revisiting Self-Training for Neural Sequence Generation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.129494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.129494Z digest=sha256:8085f661c9db8d9cbae16f70ba54dfd6dfd1ec65bf07842ee47a5367fc96a7a4

Observation 3973a647-d213-43d5-a2d9-69ab10f3e25e · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

Efficient Knowledge Injection in LLMs via Self-Distillation Quantifying Memorization Across Neural Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.099890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.099890Z digest=sha256:f2d87cc8db4033cafbeef9a0baae7c6627abf262eacf74e6a5500827afc17790

Observation 5740aedf-cc3a-4b95-bee8-5e89c618310e · outbound

This paper cites Language models are few-shot learners.

Efficient Knowledge Injection in LLMs via Self-Distillation Language models are few-shot learners

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.095907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.095907Z digest=sha256:001ca70b616c449f0fcd58bba62bdb64b01891b66415dbcfb08a22309c066287

Observation 2a09967e-9457-49e0-a6fd-06aa0687e991 · outbound

This paper cites RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture.

Efficient Knowledge Injection in LLMs via Self-Distillation RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.121350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.121350Z digest=sha256:2584434e003b5e210955a279cca74edbaf2bd206d0c5a8bae329d3bc1647894d

Observation 28f18322-0f4e-4095-974c-ca52efd91650 · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

Efficient Knowledge Injection in LLMs via Self-Distillation A General Language Assistant as a Laboratory for Alignment

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T11:52:37.091278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:52:37.091278Z digest=sha256:c84e2d9daab61dc8c8f15dd9671a2b3f78737f8411c15cf8245ac422adf863b1

Pith citing papers

Observation f90dd867-10d9-4bb4-b121-883f62cffc77 · inbound

Cartridges: Lightweight and general-purpose long context representations via self-study cites this paper.

Cartridges: Lightweight and general-purpose long context representations via self-study Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:31.647548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:31.647548Z digest=sha256:db81a3e0ded9e42a034ace17feaee948df94298db3596ba1e480db94c31da0e5

Observation e20039dd-2798-4308-8d3c-9870e2fd9dc7 · inbound

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe cites this paper.

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:26:16.201002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T17:00:49.448352Z digest=sha256:1aac22373c91d09a95079e432f1e8d25c7cabe758055a724c57d1a8766d8a0ad

Observation fa696f57-f44a-4620-86a8-feb61f88b25c · inbound

Context Memorization for Efficient Long Context Generation cites this paper.

Context Memorization for Efficient Long Context Generation Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:43:12.658948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T10:39:09.720412Z digest=sha256:89c02b130a53b3a7ad0a3b36d97d16491b84545e3b4b8986e5c93f0b8d88f416

Observation 30f0f7d0-80a4-4a18-96fd-14e643e083d5 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.803207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T10:56:13.058872Z digest=sha256:218114344b219a386cf3a7aa3cf020dfecc8c0ec9fc617f0848d821f1b5a1227

Observation 9db19840-b7a8-45e5-9367-6f811e4c40c4 · inbound

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories cites this paper.

Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-07-13T07:44:25.325808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T07:44:25.325808Z digest=sha256:cd0b81ef1b7a2fb750e37078834866802ddc06183f79dfd6d7f55cd37d634890

Observation 2cbc327f-02d1-47a7-90a5-93bd6eb5c31b · inbound

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents cites this paper.

Rethinking Continual Experience Internalization for Self-Evolving LLM Agents Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:56:47.731556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T06:29:11.398007Z digest=sha256:e6bb5516462632b208aa34b8b5683fb362365dc052fcc4c5d566703803bf5831

Observation 0d92bd82-ed75-4d1e-914f-a3c13bc959a5 · inbound

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA cites this paper.

Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA Efficient Knowledge Injection in LLMs via Self-Distillation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.802808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T05:07:58.441326Z digest=sha256:abfffd6ba238c9fb9151e7afd38b5e230c81453f8a10080d41ff71179276a515