Pith. sign in

Paper Citation Record · LEDGER

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation

As of 12 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15281 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3d4b98e-6203-4ea3-84f8-481944c31916 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.774537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.774537Z digest=sha256:775499be3c7694a19eb346b9f217016a6752fd152bc60f81ddb19a0571f91354

Observation 6e161dd6-471d-400d-a50f-9a5dba66d90e · outbound

This paper cites Fluctuation-based adaptive structured pruning for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Fluctuation-based adaptive structured pruning for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.778052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.778052Z digest=sha256:30dde750ba36e8702e82e3d5919340cdec01b9a2666734d8ad82bb3f9dd5324f

Observation 02aab24a-4e03-4b51-b3af-43518215b7cd · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.780823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.780823Z digest=sha256:a88bfc12b5e73371d85b0ba0f6f9d4f1c44120a0fa6aa9241924e1223d759cb1

Observation b49bea46-440c-4585-bf77-5a4d907f760d · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A general language assistant as a laboratory for alignment, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.783833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.783833Z digest=sha256:d0ec44d7fe40ba82e595c923944831bb507f6e4e244c0aa57ba43db3a6012b42

Observation d4aa84aa-226c-415f-94ac-a90b796e9a9c · outbound

This paper cites Mitigat- ing open-vocabulary caption hallucinations, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mitigat- ing open-vocabulary caption hallucinations, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.787009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.787009Z digest=sha256:a1f46b7e74990462d5e5bbf1f92a6028333ca536c54466aa69dcc96fdbd3cf7b

Observation bac38b26-ef54-4f65-9184-f54dbebf4816 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.790132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.790132Z digest=sha256:d7eb627fce46540e0fdaddd05c293a1b19eb6de07d0c24016a7e06145da1218e

Observation dfd730a8-8e37-4c5e-9891-b058b4c1b3cf · outbound

This paper cites Token Merging: Your ViT But Faster.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Token Merging: Your ViT But Faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.793216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.793216Z digest=sha256:4817e9a7a6597ac05aedc6dab74577225324a344b5a143afcba6feb523e12513

Observation b91a4d79-f0b6-4bf1-a0bb-ffe534159e3a · outbound

This paper cites Emerging properties in self-supervised vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Emerging properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.796341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.796341Z digest=sha256:2077ace9ad893693600273246a63474f2c03f019ac9fb8d6d7c20a0389616f15

Observation 92479429-200b-4edf-9fc0-11b80eed0654 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimization space.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Vision transformer slimming: Multi-dimension searching in continuous optimization space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.799238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.799238Z digest=sha256:d87658ac0d799c5a366e6c9b4772492c79d87cc051ebd6d031e6d3645dbc10c7

Observation b13b73c1-4e1b-42e8-b792-8ebfafee7847 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.802481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.802481Z digest=sha256:0ddade242854f23aefd74de27b2470673b517d90e148b4b4322636afdee9f840

Observation 01cc1d15-c44a-4cdf-a841-8966e90133a5 · outbound

This paper cites The lottery ticket hypothesis for pre-trained bert networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The lottery ticket hypothesis for pre-trained bert networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.698529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.805518Z digest=sha256:43ec4298b8516f428789dfc5dec8f459400eb666eb3abba7bb547051a21ed892

Observation f86d788e-745b-4e07-9c79-257d048ac7ca · outbound

This paper cites The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.689473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.808327Z digest=sha256:d6a1c08004e3cefecd6417c89b78411b9afa18821bcf97586dd00cbba612059f

Observation fee15ab5-e6a7-492a-9566-7a05fb5d132d · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A toy model of universality: Reverse engineering how networks learn group operations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.680942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.811213Z digest=sha256:01b7923efcda1b3e12645d9cf13f1b28c201ea79e6c1dbc430bacb80e98d9ba1

Observation dd75c708-4a6f-4750-a7f1-62719a01e8af · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.813816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.813816Z digest=sha256:d1b94a36d91fc97f6bae9820a686bb8790e7ba957512cc1a1ee63943aafcfd3b

Observation 3996ee05-cad3-4b1d-bff1-23af3959df71 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.816837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.816837Z digest=sha256:fc20500fe328e348eb7b028354caa9fbeccaf14e78e8ae49e1489655eabed3a2

Observation fbb301ab-1bab-4673-8ae0-54f5dd6d3740 · outbound

This paper cites Analyzing Redundancy in Pretrained Transformer Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Analyzing Redundancy in Pretrained Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.820203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.820203Z digest=sha256:6c21dc77b9ee7484b903d60b2e2540dfe555d92f0ea9081d360524aaa6d4d755

Observation 5538669b-0403-47af-876a-931014f27133 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Imagenet: A large- scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.822860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.822860Z digest=sha256:f5c603722112ac9b7a2f5e1ae6954bbab6b9f91f784d10943c4cad43a354cec2

Observation 650b0c88-dbfe-41b6-b3cb-405e0417c8d8 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Qlora: Efficient finetuning of quantized llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.825418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.825418Z digest=sha256:3e9b26e3ce69630dd14eac37f7086327b918287604d8db874a0b48b167f818d6

Observation f9609917-7f82-4be9-a255-a1b9eef0f188 · outbound

This paper cites Eventful transformers: leveraging temporal redundancy in vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Eventful transformers: leveraging temporal redundancy in vision transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.663161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.828153Z digest=sha256:669ba5b00f81b6f9637b909a4c16e1884acdbe95ba46fdbf41eb2a0458241146

Observation a1f76e72-8720-422c-a01e-9072164333cf · outbound

This paper cites A mathematical framework for transformer circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A mathematical framework for transformer circuits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.830907Z digest=sha256:b5cb3b63321ad1315a92a0e346847f46cc7f49f5374f7a9fdc60ff7be305d2cc

Observation 199df78c-ee4c-4a08-bc03-183021e11d32 · outbound

This paper cites Depgraph: Towards any structural pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Depgraph: Towards any structural pruning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.833616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.833616Z digest=sha256:1c2726fdf6ea8543ab7b0f8eb84b2ac29a4c266d6c9db158449c5079b6e481cb

Observation 4907c096-c9f4-4529-82a1-5959b9676cdd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.836416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.836416Z digest=sha256:cc4bf6f0947301ac213c1a11ea935755d087f77f7a45b1f731de3adeee9532b0

Observation 7e1c5090-53ce-408f-834f-467a84021cfa · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Transformer Feed-Forward Layers Are Key-Value Memories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.839216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.839216Z digest=sha256:2438387e888f3ba8dde6a4e586ff9eb8942f6a49bafffcb2686966fb3491570e

Observation a0ba1185-ab48-45e3-9583-0794cbcd0c7e · outbound

This paper cites Successor Heads: Recurring, Interpretable Attention Heads In The Wild.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.842513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.842513Z digest=sha256:f2b820366f542a71090e2b15072357158e3ff9511938a523c34df207aa89de50

Observation 26169097-f8ff-4db5-af4b-4befc0b14ab3 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MiniLLM: Knowledge distillation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.845580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.845580Z digest=sha256:ce26e04dea2b50dcd8ace7259cd4a9ec1b2dfbe0800897bf9948c58bcb09b395

Observation 414e3cc2-5e12-4c2a-9203-cec92ff814e8 · outbound

This paper cites Learning efficient vision transformers via fine-grained manifold distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Learning efficient vision transformers via fine-grained manifold distillation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.634244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.848506Z digest=sha256:82b87bafd60d1edbe881c3226caa268ed8863ae22a1df24fb0d80834192f407f

Observation 131feb79-a177-4bfb-a74d-01897e54cf75 · outbound

This paper cites Masked autoencoders are scalable vision learners, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Masked autoencoders are scalable vision learners, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.850924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.850924Z digest=sha256:c937115d88ef43fb98cd2d1b2e7e1e948a06b717d2e996be0e695cce799991f8

Observation 7485a9dc-058c-473a-8750-49079323c695 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation What Matters in Transformers? Not All Attention is Needed

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.853118Z digest=sha256:65925c171f9d5f5fcb783a822c2d7700d8b09e7317d76bca5cb98d5d74dd1f09

Observation eb07f159-3abb-4775-ac42-e7edb1f28aaf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.855575Z digest=sha256:57e3b998d6cadee51114de70ea8b8980f3ad1a7e5ae1082afe7f38dd818ee0a4

Observation 390edfd7-ede8-4410-9e30-cc7a3240712d · outbound

This paper cites Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:38:52.316704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.858131Z digest=sha256:7c2af9444aefe3666b7935d58a9aaa711e7a661b1c3fa869a024a93f2d638ffb

Observation ad95b5a0-f5e3-493d-a44e-54b6f12f8fb1 · outbound

This paper cites Mixture of Nested Experts: Adaptive Processing of Visual Tokens.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.860966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.860966Z digest=sha256:68b63185de681b26eecb63b95847e0bd320270ef1993a58747709e48d20dcf9f

Observation a3be0e74-894d-484d-9fa7-405a2e6d81d2 · outbound

This paper cites Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.619430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.864184Z digest=sha256:650e695279ad6713f84910f25848b6be7b691103cfd3687adbca6a055f5ebf20

Observation 55eab72c-802c-498f-82b7-411ff9c34e9a · outbound

This paper cites Expedited training of visual conditioned language generation via redundancy reduction.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Expedited training of visual conditioned language generation via redundancy reduction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.610412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.867019Z digest=sha256:ba6f1c4a00aec04a38862fe6607363b075d41616d82c0bfcadcbf6b0f3f58a31

Observation 084e9234-9b0f-4e80-addb-0417bb2a341c · outbound

This paper cites Mixtral of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.869956Z digest=sha256:e63b5c01525a958c03928fd5e74359988cdbf21d1f65ec599b2fc5e089c05db2

Observation 14383e2d-b807-48c9-b653-95ee6e53d585 · outbound

This paper cites Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit).

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.602038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.872978Z digest=sha256:9a7428cd512e62065c13278b0dfb37b9a0df7a7f9fcc1e17a9e0263d9f77ea94

Observation 682ee29e-5d31-4b4a-9497-a16b0031564e · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation TinyBERT: Distilling BERT for Natural Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.875902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.875902Z digest=sha256:e1cbeec497ce3c63f3aeebf9f71b1582815eb28da48791f2e31bae5506a4bbb5

Observation 74b67fb5-c03e-424c-8183-147344f8fe73 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.879159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.879159Z digest=sha256:17f0c76c11f4a7c4c26c175d39b140ab3caa9cbc575433154ac4a7b51137095c

Observation e73ced7e-ee60-4cf1-998e-531e43d0c297 · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-Distillation for Further Pre-training of Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.882149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.882149Z digest=sha256:ba443b6c94b07c3794c93d81f13e49b3b0b5cbf90e5fd77642569f818b02efdc

Observation 4113c3f4-3216-4e76-a6e1-48836cd0a659 · outbound

This paper cites Clustered imagenet labels for training production-friendly image classifier.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Clustered imagenet labels for training production-friendly image classifier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.594211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.884764Z digest=sha256:d801dc3a0fec17ecd2a5c807251b07e163a65885a8c78bd7d2a2579e6450b87c

Observation 5593dda4-a275-4e98-a5f3-4d9a0ef3351a · outbound

This paper cites Knowledge distillation via the target-aware transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation via the target-aware transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.887144Z digest=sha256:d1f2a417d0d0a4fd2fea69f8e4387fed5457e83874e9ac5b4e4e02153c3c7c70

Observation 8b84504f-67ac-4341-814b-58b247eef34d · outbound

This paper cites FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.889606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.889606Z digest=sha256:e889deb423bcf3f2d2b5599c68b6263a592109213ac527613f18ea486041ada6

Observation 50ac9757-c5a3-4dfb-8f42-4646693131ec · outbound

This paper cites Visual instruction tuning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.892356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.892356Z digest=sha256:e367eb71eef102cbf0f29757f4c1ead6a16cc9477f636707348314a3e0255ae4

Observation 3aa03e23-d79f-447a-8594-7510f5fce33e · outbound

This paper cites Oscillation-free quantization for low-bit vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Oscillation-free quantization for low-bit vision transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.572215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.894734Z digest=sha256:b7b40b6867a1165f8f10ecf511b00bcea4415e1eae1ce07a06a357f353ebc4eb

Observation a0cf4fec-e8a2-451e-b519-75c6bbb1ef38 · outbound

This paper cites Post-training quantization for vision transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Post-training quantization for vision transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.896973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.896973Z digest=sha256:db4452aec360e3201e1a9c5e5b93df0dc6fad05ec5800044eb5c4b3f496af575

Observation 106afdf3-ea68-44a8-a341-4635adc29ec7 · outbound

This paper cites Anytime Dense Prediction with Confidence Adaptivity.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Anytime Dense Prediction with Confidence Adaptivity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.899751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.899751Z digest=sha256:6d059d8dfe5f54cf6f0ef3f8715054d558d70e04cf2bc520248d5f434eddfd0d

Observation 7688bba1-7d21-40ad-8772-053af3629bdb · outbound

This paper cites A transformer-based model with self-distillation for multimodal emotion recognition in conversations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A transformer-based model with self-distillation for multimodal emotion recognition in conversations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.558394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.902822Z digest=sha256:9b6daaaea8eb36cebff0ac3566c45a7cfbc4380f77c9ab61c14be6bb48a7eeed

Observation 89d96134-36af-4bb5-a78c-36ab88e259ad · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Llm-pruner: On the structural pruning of large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.905570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.905570Z digest=sha256:afc8d12b10dfe0eb9484fa26479197d5f4ed21f9eb5a5cad181cfc3cdd4f10f4

Observation 8cc78494-02f8-4e58-ba88-8ed448c12c59 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Copy Suppression: Comprehensively Understanding an Attention Head

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.908556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.908556Z digest=sha256:37ad208f18270ab2fb7dc0e343241e1577a0f315e352778e7b3d20611fa3245c

Observation 0bd13a09-5670-464f-a3aa-340dc8418f03 · outbound

This paper cites Locating and editing factual associations in gpt.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Locating and editing factual associations in gpt

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.911656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.911656Z digest=sha256:b59f69b370e3dac1ecc6f4a556b8696fa7161444bfe04a89a54eaeedafd18c8c

Observation caa4206b-3f0a-453a-ae07-76d0bcfe2ca4 · outbound

This paper cites Circuit Component Reuse Across Tasks in Transformer Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Circuit Component Reuse Across Tasks in Transformer Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.914561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.914561Z digest=sha256:8f6064f6ad5c535e8c7317c91943afbfd64d10de7556e8826f282159bdabe230

Observation 2baf1a4b-320a-45e0-82ef-de20738605bb · outbound

This paper cites Zoom in: An introduction to circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Zoom in: An introduction to circuits

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.917590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.917590Z digest=sha256:fe5c90a9e2cecca140d0577c1d9c0bb7e0b8b88a948d674ac416ecdce4729d58

Observation 0ad33b7e-34c7-416e-887b-3928d987a573 · outbound

This paper cites In-context Learning and Induction Heads.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation In-context Learning and Induction Heads

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.919899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.919899Z digest=sha256:42ae42c24c8fceb0edc6d94c98c33e1772847d1658f633a19db88c99bad59039

Observation f652ae6a-d2fb-4fc5-abb2-6ba6a042452c · outbound

This paper cites Ia-red 2: Interpretability-aware redundancy reduction for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Ia-red 2: Interpretability-aware redundancy reduction for vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.534655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.922482Z digest=sha256:2dee504151a846e80a22fe30b1f464a611935bbd2376aa63d9fd3bbf78b9b0a5

Observation 5add9e11-2b4b-45b8-be8d-75586f889ba1 · outbound

This paper cites Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.526272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.924867Z digest=sha256:ee98885e2955bde44fe06dff8ed25c93cab9bada393dfb3bd315e1baca6556b6

Observation a20e40e5-ec2e-483d-b558-a50e0fc0c57f · outbound

This paper cites A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.927037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.927037Z digest=sha256:6baeeb5e294f20ea067bd43c35e800777f14acb7feae0f171dbbeff34762da35

Observation 8c018973-c0ee-4a62-9d21-bf2519e8e086 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.929776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.929776Z digest=sha256:a8503d04428467ef38349650bc1f3b34b4442aa33f2cf1558c0e1e27ba088847

Observation 0f16e29e-3e9a-41c1-9fcd-3cabd2f40520 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.932616Z digest=sha256:c4843b0d9a73ddea92326c76e3f801311890f3f8e683516b30b401d3605996e7

Observation bdf3dd24-01e7-4af5-9c2b-c8b570da2546 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.935815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.935815Z digest=sha256:a18b1a5908276a7f349df9aa62874d5ecc3ef945a0bd1bdaca4a338e2cedb49d

Observation 39d689fd-4cd5-4e07-8e20-31bc538d89af · outbound

This paper cites Confident adaptive language modeling.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Confident adaptive language modeling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.513563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.938966Z digest=sha256:6f8fdc7a0c1e5f00245f3f0b8c2eff853c3ee1afe66d42309a4ea9502653bcd0

Observation 391d1b87-f156-45bc-b39e-55f78eb52bd6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.941889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.941889Z digest=sha256:411f2ef18423dc92675a95c9176f1eb889b3fe60b3ed963d0994689359144944

Observation 9e457abb-0556-4269-abc2-bfc9f1873c4e · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.944839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.944839Z digest=sha256:8d98c6def173c75f44287f9536ce0e0420e2888ef424e89d3f8015f3d55fa38d

Observation 54e6f198-ea37-4b4c-954b-1d271ad9dd34 · outbound

This paper cites A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.947280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.947280Z digest=sha256:f2ad9170fce0cd5f621d3d35eaabc441387d6cf52473a1349c534a30d24c57e6

Observation 697491b3-5265-46f7-88e0-6c8627f8810b · outbound

This paper cites Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.505781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.949782Z digest=sha256:ff0446351c892eb96cf3185cb7a13c147ff34c04ad65bb1450e7addb273f1313

Observation 73e114a5-6a81-4076-92a5-f43ba06e4e2d · outbound

This paper cites Self-distilled vision transformer for domain generalization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distilled vision transformer for domain generalization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.496992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.952400Z digest=sha256:c90e4a92d2410d920154ce7e805d6a4a8d5c365e7d14dd611ee45386e6220ee3

Observation d3b10aa2-14bb-43f3-b59c-0b57d06a790a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.954669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.954669Z digest=sha256:2060bfa63493c99676488819942816724c76f5ad12767233b354513146bd14cf

Observation 8312c520-0e96-41d9-b045-14fd40262c2b · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.957799Z digest=sha256:8aef01abeb9f0cb8c1cca9d64450f9d32d5d7be6712348bf1b4b4c34bef2c7b5

Observation b47718a4-6be5-4669-90a5-67c84787c257 · outbound

This paper cites Patch slimming for efficient vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patch slimming for efficient vision transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.960881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.960881Z digest=sha256:e1af1f7a52ee4c3b64e1061582cf9a86ee7e465ebe924aceaed27718ed92604f

Observation fc3cce1a-9e23-42a2-8fba-d94c5de0578b · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.483135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.963709Z digest=sha256:921160409c41cd18eadb8e21fc1e39f80eb22e98c37a893c59144b42fbf6a33a

Observation 8378a85f-e41e-4276-9c62-43ec2ac1af75 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training data-efficient image transformers & distillation through attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.966919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.966919Z digest=sha256:9fc8a9ce9da7b90ce2eabde1ffc00a9c4188463f2cac6779fa6af0cefa2553d6

Observation e3fa46d1-7012-4a75-910a-88ec06df7c40 · outbound

This paper cites Attention is all you need.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.969832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.969832Z digest=sha256:d3d7ca1e038cde48ba2029cb7bffb0539fb8a6d70f49baac47961a93613e940a

Observation 03c70042-9d43-43c0-8c76-c5f78c24529f · outbound

This paper cites Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.464845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.972731Z digest=sha256:172bc7a78da88a3834c687e7ef5977ffb3df0020fbd56085ebbb6226734ea71a

Observation 8adb33bb-3086-440e-b335-7cab20402372 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.975653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.975653Z digest=sha256:5628f177f5096a620321864498c7b9dc064ca2f57d0073412299374b1c740495

Observation b62e5c21-9f43-4b81-ab10-ab17702fb178 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.456906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.979027Z digest=sha256:5ff150e066dc134920490e51847937b14847706805cfb7c70ae170659f2d8bda

Observation e24b8f23-3957-4d42-9397-6e0ec64b77d3 · outbound

This paper cites Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.448430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.982279Z digest=sha256:5b035a9ab7b13bf52e7e1c3349cbabdb5ac0922dc5a0e749e6b2ce584692dfef

Observation 1749f87a-2a79-4e35-ba69-ae658f8d1b8d · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.438892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.984954Z digest=sha256:888b3d5f68149c858d2eed9351dc3165f85a0e3bdf48f1d6392d0466f693643b

Observation 23cc25b4-72e7-4e58-bf94-be7ee936f505 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.987271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.987271Z digest=sha256:47da403618b5e3b1e9fec39afffcf1c00468aa50ec442e1e6d260c15b16dda6c

Observation 233e4718-6d55-4085-a644-6d508d92c1f0 · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Structured Pruning Learns Compact and Accurate Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.990020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.990020Z digest=sha256:80abb46334f2489f2659b7794ed6ff4086287560f170ebf3364c3d9ee21fa3b0

Observation 33c2b9fc-d032-40a7-95a7-0384e79a0ba4 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.992893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.992893Z digest=sha256:9d98886a1d9a5360b348a8f37962b925ecf03905a9ca987e4d1d082331322699

Observation 33c4be17-a40d-4aa6-935b-01e311898b3d · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.995190Z digest=sha256:c29065f72d37bd51854e84dfa934533a59e539cafa98f56858f10cd3fc5b8687

Observation 0273513d-35ee-4930-a305-18c2edda9540 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation X-pruner: explainable pruning for vision transformers

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.424277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:51.998085Z digest=sha256:01a06aa11e0cd61e55070ff942db61dbf3bd4f24225205f0e2bf857d02caa818

Observation 93c0355e-f848-45db-9b2f-12e6afd37318 · outbound

This paper cites Unified Visual Transformer Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unified Visual Transformer Compression

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.000864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.000864Z digest=sha256:925447e6fe19e485b5a470978d6200a95a7b0c6eb3ac1fe2db35174966be2b60

Observation 11a5dcf9-149d-4f7b-b713-3b8c2f8a1686 · outbound

This paper cites MoEfication: Transformer Feed-forward Layers are Mixtures of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.004555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.004555Z digest=sha256:f3adc79d16360fc296903dcf1c7c6712edfea9bc78fea570b23832d6c768483b

Observation caea50ba-6c6e-424d-bb5c-0cc1ee2e5027 · outbound

This paper cites Knowledge distillation based on transformed teacher matching.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation based on transformed teacher matching

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.414808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:52.007595Z digest=sha256:a7afb1a8836aca6c0512f30e9c569e67e533ad54824326096520d109741ff743

Observation 80958d0c-b3c7-4749-8de3-4432ec47e4ca · outbound

This paper cites The clock and the pizza: Two stories in mechanistic explanation of neural networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The clock and the pizza: Two stories in mechanistic explanation of neural networks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.010319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.010319Z digest=sha256:94ed62c846e317fecde6f11001f573812e30b2b85236fb9f93f5ac65af737909

Observation 19db123c-225d-4678-83db-042e11c80ad4 · outbound

This paper cites LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.013139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.013139Z digest=sha256:f10138998f9376bd011a644383f4da131a9d44151d31664f4797f6961f7d067b

Observation e89b4d73-6cb9-4692-bfff-0be0ab6974aa · outbound

This paper cites MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.016407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.016407Z digest=sha256:10855d0ee06650ab6459b14869128cabb155c703211847f4f22557425c4806a3

Observation 2a226cfb-df33-4745-a251-066a3616abc2 · outbound

This paper cites Possible solutions could include:.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Possible solutions could include:

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.402343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:52.019705Z digest=sha256:48d0b035c6516918860265ea359043babff17ef8f0aa62ce1e4f916f781a0c04

Observation da8cde43-a337-4086-a708-a2e5e522a007 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.394336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:52.022253Z digest=sha256:10f7c5f7d9ffdcb4b849e3f924f958e275147e3b5890a22ee6b7eefa7b6a9027

Observation 35c4f703-c447-4c33-ad9e-f9b61206eaa1 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.385463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T14:38:52.024717Z digest=sha256:c59a802afc1b00087d2454f22a4df48dd0130d57689cd1a802af06af5533841e

Pith citing papers

No inbound Pith citation observations are available.