Pith. sign in

Paper Citation Record · LEDGER

Patient Knowledge Distillation for BERT Model Compression

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:1908.09355.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.09355 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:57:25.814259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T21:17:35.850483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b50a58fd-6a2e-4d9d-b589-059f7575d5b6 · inbound

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations cites this paper.

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations Patient Knowledge Distillation for BERT Model Compression

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:26:58.120212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T12:26:58.015594Z digest=sha256:cf5d890745689541c9743349ccb281af1d9054d2c8a9669d2c8fa2089c54882e

Observation a64d7f3e-7630-4f18-9fc8-1525884bf3e0 · inbound

Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed cites this paper.

Knowledge Distillation in Iterative Generative Models for Improved Sampling Speed Patient Knowledge Distillation for BERT Model Compression

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:56:19.902215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T03:56:19.837017Z digest=sha256:f7e29819f5bc969c5c8a766c8183343e2f173cf601b3a1d7e7651a778bc40baa

Observation 53667ac5-4a18-447c-846b-fd3c0643f1c1 · inbound

BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as Experts cites this paper.

BEEM: Boosting Performance of Early Exit DNNs using Multi-Exit Classifiers as Experts Patient Knowledge Distillation for BERT Model Compression

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T17:57:25.814259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:57:25.814259Z digest=sha256:3b4d735b790be7f98083db6868d39ac5bcbf774ca329884418f4ce29b2fd48be

Observation 1dc69456-7829-440e-b0f6-71fb2c3370e2 · inbound

MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual Learning cites this paper.

MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual Learning Patient Knowledge Distillation for BERT Model Compression

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T12:26:25.131696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:26:25.131696Z digest=sha256:547621fa86684be68e903de9bb44651ffa4d3b1babc90a53f6af6390b4a237ce

Observation 3a7de572-72d7-4a51-80b2-fa93d0779bc4 · inbound

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers cites this paper.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Patient Knowledge Distillation for BERT Model Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.725290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.725290Z digest=sha256:4e278c8835d09c323378f3b98478081b4ea41c58adbe620daa012fa5d8096721

Observation 9ba242ea-e607-45c4-8656-c21c32ea2bf0 · inbound

Olica: Efficient Structured Pruning of Large Language Models without Retraining cites this paper.

Olica: Efficient Structured Pruning of Large Language Models without Retraining Patient Knowledge Distillation for BERT Model Compression

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:19:39.455568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:19:39.455568Z digest=sha256:8d9a1f9bdf7392ea70c2f82fc416ca9978898cf0dfd7e02c4086c60744ed0b6e

Observation 9b38405d-135d-4eb8-af14-40f04c586881 · inbound

AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes cites this paper.

AgentDistill: Training-Free Agent Distillation with Generalizable MCP Boxes Patient Knowledge Distillation for BERT Model Compression

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:34.333057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:34.333057Z digest=sha256:f8ca0a771b7dd646771ede1a7aad998db1e19c2ac730cba5661d5a92c7b59b4d

Observation 6672605b-9dc7-43f7-83da-6d37ccf146d8 · inbound

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models cites this paper.

GenRecal: Generation after Recalibration from Large to Small Vision-Language Models Patient Knowledge Distillation for BERT Model Compression

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:25.971549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:25.971549Z digest=sha256:37ae1abeedbc1e3dc82e286062ffdb52142c70813e2c4fbcecc6256c4e539253

Observation 6ad1efc8-5f6b-46c8-a46a-d5854f7f1441 · inbound

MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants cites this paper.

MiCoTA: Bridging the Learnability Gap with Intermediate CoT and Teacher Assistants Patient Knowledge Distillation for BERT Model Compression

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:22.171701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:22.171701Z digest=sha256:266bca76ee7afebb5d3be9b07d039378767943c3a7afd532832477cd9b749306

Observation 427cc078-087e-4d6b-a1d0-186875127e40 · inbound

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study cites this paper.

Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study Patient Knowledge Distillation for BERT Model Compression

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:47.217308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:47.217308Z digest=sha256:3f6ff77a7bffa5e4d32c693064528f36289e12e0f005ba06cb4efa49fe5c8739

Observation da3e02a0-b92f-4709-a7c2-6d2542017fcf · inbound

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code cites this paper.

Model Compression vs. Adversarial Robustness: An Empirical Study on Language Models for Code Patient Knowledge Distillation for BERT Model Compression

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:01:55.914497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:01:42.228190Z digest=sha256:f22ff806a9ccce5933a97617326936768302452c9125f7e1765f6f7ec59fc777

Observation b9e07d11-0b8e-46a7-b308-b60e3ba51776 · inbound

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization cites this paper.

TinyMusician: On-Device Music Generation with Knowledge Distillation and Mixed Precision Quantization Patient Knowledge Distillation for BERT Model Compression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T13:12:14.557846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:12:14.557846Z digest=sha256:821b9fdf4cc637df594ccdbc352b1f48481db16543428aa1dd733068656161ae

Observation 59a2f1f3-0e4b-414a-9f2c-7b1e7f9e0058 · inbound

A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher? cites this paper.

A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher? Patient Knowledge Distillation for BERT Model Compression

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:50:31.924965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T23:45:54.490984Z digest=sha256:b1f2d9b096eba3dac0e625b6a9e9197bc9d75e869eee589e68e5c4a08507f565

Observation 49097048-3337-4c14-a86f-4948f7eaf24e · inbound

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency cites this paper.

Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency Patient Knowledge Distillation for BERT Model Compression

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:56:14.132583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:47:38.100037Z digest=sha256:ceeb18b7e1753529cc106c999fe3a24d95f9af3d8f27bfdad28f0d9bbb9b548e

Observation 44dd3aee-e895-49b9-a19f-10385be2c025 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:46:42.823215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:898e8d033f99381b0522e771a59c70f6ef80a676abd3ec5ec4a703ed870ff83e

Observation bf0a0a9f-1388-4073-81df-6711eb51fc0f · inbound

Curriculum Learning-Guided Progressive Distillation in Large Language Models cites this paper.

Curriculum Learning-Guided Progressive Distillation in Large Language Models Patient Knowledge Distillation for BERT Model Compression

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:05.779234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:53:21.920498Z digest=sha256:d60c47b02a20db3b1bd235389b21b929d4c49135d685484595ae4fd6b5a89d4e

Observation 1a3b452b-9c5c-4503-8765-bffb1dc09500 · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:24:39.185969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T12:23:42.587689Z digest=sha256:ca455ffc996153b7a438c8bbd96b0181db7fc10b1721e495b0c6ccde652dcb2b

Observation 53bb5af7-8338-4776-b903-b170b661013b · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.229372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-04T00:38:33.708836Z digest=sha256:536781b3a57d5c792ea5b8e38d34e0cb68112b243351c34efae0c70513ff61d6

Observation d7be940a-e20c-4407-add2-74746e405e6a · inbound

NITP: Next Implicit Token Prediction for LLM Pre-training cites this paper.

NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d3f82926589551cb459855ddcb03184bfe0cd7ea43d69296553d0657c394d3c8

Observation c050cf69-5935-4071-9ffd-fdf07422f14b · inbound

Toward Calibrated, Fair, and accurate Deepfake Detection cites this paper.

Toward Calibrated, Fair, and accurate Deepfake Detection Patient Knowledge Distillation for BERT Model Compression

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-28T07:11:45.307920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T07:05:18.026601Z digest=sha256:f14be9118f27ee084521e67941f410432498566d596fad8b9f509eb4e2762157

Observation da0caf49-bae3-492c-aa3f-4097e546ee0b · inbound

Scaling Laws for Task-Specific LLM Distillation cites this paper.

Scaling Laws for Task-Specific LLM Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:40:00.673237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:29:49.787477Z digest=sha256:4c7e97dc4c7af829f006cd921f84126e1d46d2444520d09e0a25d10a64b2cf68

Observation f6820243-934d-4162-b0aa-614faeb69a46 · inbound

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG cites this paper.

Distill Where the Student Goes: Teacher-Regularized RL for English-Evidence Cross-Lingual RAG Patient Knowledge Distillation for BERT Model Compression

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T05:47:14.970921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:47:14.970921Z digest=sha256:2ce2ea304bc78f9297e6e5404d86b2f67818828eb0e53a00e6a53d3417decaf9

Observation 55f7a977-46e9-4096-a08e-95778e262110 · inbound

Enhancing deep learning models for time series classification via knowledge distillation cites this paper.

Enhancing deep learning models for time series classification via knowledge distillation Patient Knowledge Distillation for BERT Model Compression

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-10T21:17:35.864696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T21:09:54.557270Z digest=sha256:cc24540aa5d83957b4b965f7b737effc216688e87ad3104eb4987e3faf8592d9

Observation fd44600d-a514-43ac-ab1d-bc1b4840391a · inbound

Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression cites this paper.

Progressive$^2$: A Teacher-Student Progressive Co-Evolving Knowledge Distillation Method for Substantial Model Compression Patient Knowledge Distillation for BERT Model Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:43.093855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:41:43.093855Z digest=sha256:e3844a4e32bef03d9aa1e435d339c5b1fae4b9098147169433f1da313afc2b0d

Observation 7326ea17-1aee-4884-8c93-faf489794710 · inbound

BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning cites this paper.

BnBERT-iPET: Sparse Few-Shot Language Modeling for Bengali via Lottery Ticket Pruning Patient Knowledge Distillation for BERT Model Compression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T05:20:01.478142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:20:01.478142Z digest=sha256:62c9e7b741fcd30d88e8327c7076f26cbdb7f32f417acce36a5cab615cdc228e