Pith. sign in

Paper Citation Record · LEDGER

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation

As of 13 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15281.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15281 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved64
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3d4b98e-6203-4ea3-84f8-481944c31916 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.774537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.774537Z digest=sha256:775499be3c7694a19eb346b9f217016a6752fd152bc60f81ddb19a0571f91354

Observation 6e161dd6-471d-400d-a50f-9a5dba66d90e · outbound

This paper cites Fluctuation-based adaptive structured pruning for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Fluctuation-based adaptive structured pruning for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.778052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.778052Z digest=sha256:30dde750ba36e8702e82e3d5919340cdec01b9a2666734d8ad82bb3f9dd5324f

Observation 02aab24a-4e03-4b51-b3af-43518215b7cd · outbound

This paper cites Dynamic context pruning for efficient and interpretable autoregressive transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamic context pruning for efficient and interpretable autoregressive transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.780823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.780823Z digest=sha256:a88bfc12b5e73371d85b0ba0f6f9d4f1c44120a0fa6aa9241924e1223d759cb1

Observation b49bea46-440c-4585-bf77-5a4d907f760d · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A general language assistant as a laboratory for alignment, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.783833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.783833Z digest=sha256:d0ec44d7fe40ba82e595c923944831bb507f6e4e244c0aa57ba43db3a6012b42

Observation d4aa84aa-226c-415f-94ac-a90b796e9a9c · outbound

This paper cites Mitigat- ing open-vocabulary caption hallucinations, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mitigat- ing open-vocabulary caption hallucinations, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.787009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.787009Z digest=sha256:a1f46b7e74990462d5e5bbf1f92a6028333ca536c54466aa69dcc96fdbd3cf7b

Observation bac38b26-ef54-4f65-9184-f54dbebf4816 · outbound

This paper cites GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.790132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.790132Z digest=sha256:d7eb627fce46540e0fdaddd05c293a1b19eb6de07d0c24016a7e06145da1218e

Observation dfd730a8-8e37-4c5e-9891-b058b4c1b3cf · outbound

This paper cites Token Merging: Your ViT But Faster.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Token Merging: Your ViT But Faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.793216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.793216Z digest=sha256:33110527a478dc9136c5038e26adfff4535174e35dfb5140fb2d9e287a2b0764

Observation b91a4d79-f0b6-4bf1-a0bb-ffe534159e3a · outbound

This paper cites Emerging properties in self-supervised vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Emerging properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.796341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.796341Z digest=sha256:2077ace9ad893693600273246a63474f2c03f019ac9fb8d6d7c20a0389616f15

Observation 92479429-200b-4edf-9fc0-11b80eed0654 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimization space.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Vision transformer slimming: Multi-dimension searching in continuous optimization space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.799238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.799238Z digest=sha256:d87658ac0d799c5a366e6c9b4772492c79d87cc051ebd6d031e6d3645dbc10c7

Observation b13b73c1-4e1b-42e8-b792-8ebfafee7847 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.802481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.802481Z digest=sha256:0ddade242854f23aefd74de27b2470673b517d90e148b4b4322636afdee9f840

Observation 01cc1d15-c44a-4cdf-a841-8966e90133a5 · outbound

This paper cites The lottery ticket hypothesis for pre-trained bert networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The lottery ticket hypothesis for pre-trained bert networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.698529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.805518Z digest=sha256:e09ac4af16d8c4babba1444f8306bb1951c73ef717b5e5510efbb7839a4a0b24

Observation f86d788e-745b-4e07-9c79-257d048ac7ca · outbound

This paper cites The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.689473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.808327Z digest=sha256:d1c21014ad8554fa9b6905dfe5cb785c2a7add752a757cdda18560494e4ba45e

Observation fee15ab5-e6a7-492a-9566-7a05fb5d132d · outbound

This paper cites A toy model of universality: Reverse engineering how networks learn group operations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A toy model of universality: Reverse engineering how networks learn group operations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.680942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.811213Z digest=sha256:10c7c7202dc1c88e117f3243e1abfddb1d0f36de35da24459bdc6256d8a58fb7

Observation dd75c708-4a6f-4750-a7f1-62719a01e8af · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.813816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.813816Z digest=sha256:d1b94a36d91fc97f6bae9820a686bb8790e7ba957512cc1a1ee63943aafcfd3b

Observation 3996ee05-cad3-4b1d-bff1-23af3959df71 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.816837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.816837Z digest=sha256:3116bdb54ef17db0f2fa82b0f44c960960ec336391938305260a7619e3339803

Observation fbb301ab-1bab-4673-8ae0-54f5dd6d3740 · outbound

This paper cites Analyzing Redundancy in Pretrained Transformer Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Analyzing Redundancy in Pretrained Transformer Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.820203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.820203Z digest=sha256:6c21dc77b9ee7484b903d60b2e2540dfe555d92f0ea9081d360524aaa6d4d755

Observation 5538669b-0403-47af-876a-931014f27133 · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Imagenet: A large- scale hierarchical image database

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.822860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.822860Z digest=sha256:f5c603722112ac9b7a2f5e1ae6954bbab6b9f91f784d10943c4cad43a354cec2

Observation 650b0c88-dbfe-41b6-b3cb-405e0417c8d8 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Qlora: Efficient finetuning of quantized llms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.825418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.825418Z digest=sha256:3e9b26e3ce69630dd14eac37f7086327b918287604d8db874a0b48b167f818d6

Observation f9609917-7f82-4be9-a255-a1b9eef0f188 · outbound

This paper cites Eventful transformers: leveraging temporal redundancy in vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Eventful transformers: leveraging temporal redundancy in vision transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.663161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.828153Z digest=sha256:dcbc06e36cc865af3af0986047755aaacd401df3b26cb962a6a60decbe31cb65

Observation a1f76e72-8720-422c-a01e-9072164333cf · outbound

This paper cites A mathematical framework for transformer circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A mathematical framework for transformer circuits

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.830907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.830907Z digest=sha256:b5cb3b63321ad1315a92a0e346847f46cc7f49f5374f7a9fdc60ff7be305d2cc

Observation 199df78c-ee4c-4a08-bc03-183021e11d32 · outbound

This paper cites Depgraph: Towards any structural pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Depgraph: Towards any structural pruning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.833616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.833616Z digest=sha256:1c2726fdf6ea8543ab7b0f8eb84b2ac29a4c266d6c9db158449c5079b6e481cb

Observation 4907c096-c9f4-4529-82a1-5959b9676cdd · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.836416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.836416Z digest=sha256:cc4bf6f0947301ac213c1a11ea935755d087f77f7a45b1f731de3adeee9532b0

Observation 7e1c5090-53ce-408f-834f-467a84021cfa · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Transformer Feed-Forward Layers Are Key-Value Memories

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.839216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.839216Z digest=sha256:b5b6283e3a7b3cb3fc9c99a2e9fad1612d22c5a6d3f20844db42b44342e3083a

Observation a0ba1185-ab48-45e3-9583-0794cbcd0c7e · outbound

This paper cites Successor Heads: Recurring, Interpretable Attention Heads In The Wild.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Successor Heads: Recurring, Interpretable Attention Heads In The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.842513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.842513Z digest=sha256:0549c51eed7665d1d11402025fccd3220ebe87e1ee8a464f8f0f35c0e2a87f50

Observation 26169097-f8ff-4db5-af4b-4befc0b14ab3 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MiniLLM: Knowledge distillation of large language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.845580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.845580Z digest=sha256:ce26e04dea2b50dcd8ace7259cd4a9ec1b2dfbe0800897bf9948c58bcb09b395

Observation 414e3cc2-5e12-4c2a-9203-cec92ff814e8 · outbound

This paper cites Learning efficient vision transformers via fine-grained manifold distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Learning efficient vision transformers via fine-grained manifold distillation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.634244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.848506Z digest=sha256:ed2157f51a33393bd02e7c538cc2647b74157b5efbc3340226679489a1a1808d

Observation 131feb79-a177-4bfb-a74d-01897e54cf75 · outbound

This paper cites Masked autoencoders are scalable vision learners, 2021.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Masked autoencoders are scalable vision learners, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.850924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.850924Z digest=sha256:c937115d88ef43fb98cd2d1b2e7e1e948a06b717d2e996be0e695cce799991f8

Observation 7485a9dc-058c-473a-8750-49079323c695 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation What Matters in Transformers? Not All Attention is Needed

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.853118Z digest=sha256:4cc0c7d230db179b1fa8e8f3cfa3768c0fbc9bd70f07e1a645c2aa5dca171d0e

Observation eb07f159-3abb-4775-ac42-e7edb1f28aaf · outbound

This paper cites Distilling the Knowledge in a Neural Network.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Distilling the Knowledge in a Neural Network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.855575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.855575Z digest=sha256:57e3b998d6cadee51114de70ea8b8980f3ad1a7e5ae1082afe7f38dd818ee0a4

Observation 390edfd7-ede8-4410-9e30-cc7a3240712d · outbound

This paper cites Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:38:52.316704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.858131Z digest=sha256:a50a6de4e9d4f0532b5d225d05659126954ed109f0502bdfdab30f563156ee4c

Observation ad95b5a0-f5e3-493d-a44e-54b6f12f8fb1 · outbound

This paper cites Mixture of Nested Experts: Adaptive Processing of Visual Tokens.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.860966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.860966Z digest=sha256:68b63185de681b26eecb63b95847e0bd320270ef1993a58747709e48d20dcf9f

Observation a3be0e74-894d-484d-9fa7-405a2e6d81d2 · outbound

This paper cites Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.619430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.864184Z digest=sha256:b1a4457248039a46c6eade0a06f996a15d70091db1f7b0a3b2ca83637e5d5b6d

Observation 55eab72c-802c-498f-82b7-411ff9c34e9a · outbound

This paper cites Expedited training of visual conditioned language generation via redundancy reduction.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Expedited training of visual conditioned language generation via redundancy reduction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.610412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.867019Z digest=sha256:5ec30524d275e641607c63c15b2ad8a62327d26a604178c14bb789461f70d7db

Observation 084e9234-9b0f-4e80-addb-0417bb2a341c · outbound

This paper cites Mixtral of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.869956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.869956Z digest=sha256:e63b5c01525a958c03928fd5e74359988cdbf21d1f65ec599b2fc5e089c05db2

Observation 14383e2d-b807-48c9-b653-95ee6e53d585 · outbound

This paper cites Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit).

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.602038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.872978Z digest=sha256:dd070fc4c1934f10698be6010e1c16440ae34126f540826538d5950444f22e44

Observation 682ee29e-5d31-4b4a-9497-a16b0031564e · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation TinyBERT: Distilling BERT for Natural Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.875902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.875902Z digest=sha256:e1cbeec497ce3c63f3aeebf9f71b1582815eb28da48791f2e31bae5506a4bbb5

Observation 74b67fb5-c03e-424c-8183-147344f8fe73 · outbound

This paper cites Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.879159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.879159Z digest=sha256:cf505b037a7bd26f47158e68760426c97ee46942d3e967bd04d4f680e9f60418

Observation e73ced7e-ee60-4cf1-998e-531e43d0c297 · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-Distillation for Further Pre-training of Transformers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.882149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.882149Z digest=sha256:ba443b6c94b07c3794c93d81f13e49b3b0b5cbf90e5fd77642569f818b02efdc

Observation 4113c3f4-3216-4e76-a6e1-48836cd0a659 · outbound

This paper cites Clustered imagenet labels for training production-friendly image classifier.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Clustered imagenet labels for training production-friendly image classifier

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.594211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.884764Z digest=sha256:448d547b7171b68ce4c986bc9c591654f70057bbf6c2fc75f0751fea1ef9cb48

Observation 5593dda4-a275-4e98-a5f3-4d9a0ef3351a · outbound

This paper cites Knowledge distillation via the target-aware transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation via the target-aware transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.887144Z digest=sha256:0214d66e1e9a77e82a5cbe0bef73cecb2ac6b6c051f11ea4312daf53ad3aa70d

Observation 8b84504f-67ac-4341-814b-58b247eef34d · outbound

This paper cites FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.889606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.889606Z digest=sha256:e889deb423bcf3f2d2b5599c68b6263a592109213ac527613f18ea486041ada6

Observation 50ac9757-c5a3-4dfb-8f42-4646693131ec · outbound

This paper cites Visual instruction tuning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Visual instruction tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.892356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.892356Z digest=sha256:e367eb71eef102cbf0f29757f4c1ead6a16cc9477f636707348314a3e0255ae4

Observation 3aa03e23-d79f-447a-8594-7510f5fce33e · outbound

This paper cites Oscillation-free quantization for low-bit vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Oscillation-free quantization for low-bit vision transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.572215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.894734Z digest=sha256:c6439825a25db1eb3b36de0cb739562c0b6b266b4eb7c3cc3ace98338c48f016

Observation a0cf4fec-e8a2-451e-b519-75c6bbb1ef38 · outbound

This paper cites Post-training quantization for vision transformer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Post-training quantization for vision transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.896973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.896973Z digest=sha256:db4452aec360e3201e1a9c5e5b93df0dc6fad05ec5800044eb5c4b3f496af575

Observation 106afdf3-ea68-44a8-a341-4635adc29ec7 · outbound

This paper cites Anytime Dense Prediction with Confidence Adaptivity.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Anytime Dense Prediction with Confidence Adaptivity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.899751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.899751Z digest=sha256:6d059d8dfe5f54cf6f0ef3f8715054d558d70e04cf2bc520248d5f434eddfd0d

Observation 7688bba1-7d21-40ad-8772-053af3629bdb · outbound

This paper cites A transformer-based model with self-distillation for multimodal emotion recognition in conversations.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A transformer-based model with self-distillation for multimodal emotion recognition in conversations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.558394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.902822Z digest=sha256:d618cfe7259d98db5c79ad64bf189c32f53119b2e45a215cb75230467f231ae7

Observation 89d96134-36af-4bb5-a78c-36ab88e259ad · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Llm-pruner: On the structural pruning of large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.905570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.905570Z digest=sha256:afc8d12b10dfe0eb9484fa26479197d5f4ed21f9eb5a5cad181cfc3cdd4f10f4

Observation 8cc78494-02f8-4e58-ba88-8ed448c12c59 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Copy Suppression: Comprehensively Understanding an Attention Head

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.908556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.908556Z digest=sha256:27fd648bf05a539285ce82a9c5e0e6b7880f5de629b89e40e102ed0c58bfa7c6

Observation 0bd13a09-5670-464f-a3aa-340dc8418f03 · outbound

This paper cites Locating and editing factual associations in gpt.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Locating and editing factual associations in gpt

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.911656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.911656Z digest=sha256:b59f69b370e3dac1ecc6f4a556b8696fa7161444bfe04a89a54eaeedafd18c8c

Observation caa4206b-3f0a-453a-ae07-76d0bcfe2ca4 · outbound

This paper cites Circuit Component Reuse Across Tasks in Transformer Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Circuit Component Reuse Across Tasks in Transformer Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.914561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.914561Z digest=sha256:590aa6322a74774d284312a00155b92e871739ae8b1cee9645c32285eea76e1f

Observation 2baf1a4b-320a-45e0-82ef-de20738605bb · outbound

This paper cites Zoom in: An introduction to circuits.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Zoom in: An introduction to circuits

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.917590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.917590Z digest=sha256:fe5c90a9e2cecca140d0577c1d9c0bb7e0b8b88a948d674ac416ecdce4729d58

Observation 0ad33b7e-34c7-416e-887b-3928d987a573 · outbound

This paper cites In-context Learning and Induction Heads.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation In-context Learning and Induction Heads

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.919899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.919899Z digest=sha256:42ae42c24c8fceb0edc6d94c98c33e1772847d1658f633a19db88c99bad59039

Observation f652ae6a-d2fb-4fc5-abb2-6ba6a042452c · outbound

This paper cites Ia-red 2: Interpretability-aware redundancy reduction for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Ia-red 2: Interpretability-aware redundancy reduction for vision transformers

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.534655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.922482Z digest=sha256:30e8d38721011179ca0f85b2ad7f41cef80cc43dfb4529adcd53d109e4821185

Observation 5add9e11-2b4b-45b8-be8d-75586f889ba1 · outbound

This paper cites Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.526272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.924867Z digest=sha256:c1f5983c8902782ba5f0806f514d693ad421126a4a4fec464067d264a29b67da

Observation a20e40e5-ec2e-483d-b558-a50e0fc0c57f · outbound

This paper cites A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.927037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.927037Z digest=sha256:6baeeb5e294f20ea067bd43c35e800777f14acb7feae0f171dbbeff34762da35

Observation 8c018973-c0ee-4a62-9d21-bf2519e8e086 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.929776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.929776Z digest=sha256:a8503d04428467ef38349650bc1f3b34b4442aa33f2cf1558c0e1e27ba088847

Observation 0f16e29e-3e9a-41c1-9fcd-3cabd2f40520 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.932616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.932616Z digest=sha256:c4843b0d9a73ddea92326c76e3f801311890f3f8e683516b30b401d3605996e7

Observation bdf3dd24-01e7-4af5-9c2b-c8b570da2546 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.935815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.935815Z digest=sha256:a18b1a5908276a7f349df9aa62874d5ecc3ef945a0bd1bdaca4a338e2cedb49d

Observation 39d689fd-4cd5-4e07-8e20-31bc538d89af · outbound

This paper cites Confident adaptive language modeling.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Confident adaptive language modeling

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.513563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.938966Z digest=sha256:b4e315c8d52ada90c847e1948193cabf5a64ffac93850f8483df841358d95361

Observation 391d1b87-f156-45bc-b39e-55f78eb52bd6 · outbound

This paper cites OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.941889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.941889Z digest=sha256:499fe4296c4e9020ea1f52293e690bd4f29d987ca3e43b7a2bbdfce556ccc8eb

Observation 9e457abb-0556-4269-abc2-bfc9f1873c4e · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.944839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.944839Z digest=sha256:678caf39388c5ac02f9c2150ea9f99ffd290044e46def5c7bef91488d05179a6

Observation 54e6f198-ea37-4b4c-954b-1d271ad9dd34 · outbound

This paper cites A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.947280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.947280Z digest=sha256:a8661686b95092cf09bbfe64032d77539735245951944d58a8e14bcc74413efb

Observation 697491b3-5265-46f7-88e0-6c8627f8810b · outbound

This paper cites Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.505781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.949782Z digest=sha256:c4d34a162932c9133bf799221b70be95812d86a8fcaa7ccb98f3812840d6e048

Observation 73e114a5-6a81-4076-92a5-f43ba06e4e2d · outbound

This paper cites Self-distilled vision transformer for domain generalization.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distilled vision transformer for domain generalization

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.496992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.952400Z digest=sha256:e334488d38a89cb8c31dd751bc404971c3edfca2adfa37630e396d0699bbed4b

Observation d3b10aa2-14bb-43f3-b59c-0b57d06a790a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patient Knowledge Distillation for BERT Model Compression

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.954669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.954669Z digest=sha256:2060bfa63493c99676488819942816724c76f5ad12767233b354513146bd14cf

Observation 8312c520-0e96-41d9-b045-14fd40262c2b · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.957799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.957799Z digest=sha256:8aef01abeb9f0cb8c1cca9d64450f9d32d5d7be6712348bf1b4b4c34bef2c7b5

Observation b47718a4-6be5-4669-90a5-67c84787c257 · outbound

This paper cites Patch slimming for efficient vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patch slimming for efficient vision transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.960881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.960881Z digest=sha256:e1af1f7a52ee4c3b64e1061582cf9a86ee7e465ebe924aceaed27718ed92604f

Observation fc3cce1a-9e23-42a2-8fba-d94c5de0578b · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.483135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.963709Z digest=sha256:c29b70ed50a9264ba08b00f31d102e2e66627a7f0fd13dbbeebd6c65c2f3972e

Observation 8378a85f-e41e-4276-9c62-43ec2ac1af75 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training data-efficient image transformers & distillation through attention

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.966919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.966919Z digest=sha256:9fc8a9ce9da7b90ce2eabde1ffc00a9c4188463f2cac6779fa6af0cefa2553d6

Observation e3fa46d1-7012-4a75-910a-88ec06df7c40 · outbound

This paper cites Attention is all you need.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Attention is all you need

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.969832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.969832Z digest=sha256:d3d7ca1e038cde48ba2029cb7bffb0539fb8a6d70f49baac47961a93613e940a

Observation 03c70042-9d43-43c0-8c76-c5f78c24529f · outbound

This paper cites Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.464845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.972731Z digest=sha256:c000fabb9aa415f953f4ef74572dda2d97ac98669522c1ea94a11aa96c4c7733

Observation 8adb33bb-3086-440e-b335-7cab20402372 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.975653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.975653Z digest=sha256:5628f177f5096a620321864498c7b9dc064ca2f57d0073412299374b1c740495

Observation b62e5c21-9f43-4b81-ab10-ab17702fb178 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.456906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.979027Z digest=sha256:83c8d5cfeb0e052787ad2af931a6d10bfdd4ca9452adb76eefa5d4bd9e43dcea

Observation e24b8f23-3957-4d42-9397-6e0ec64b77d3 · outbound

This paper cites Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.448430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.982279Z digest=sha256:a1479abe6020e3a8f31eb9131ed1954ac050e48f6eaa6e9bfddfe1706ea43481

Observation 1749f87a-2a79-4e35-ba69-ae658f8d1b8d · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tinyvit: Fast pretraining distillation for small vision transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.438892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.984954Z digest=sha256:316338b44f67a3d253605e694af4a36842ec207c33938b27ae64ae06639e7b33

Observation 23cc25b4-72e7-4e58-bf94-be7ee936f505 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.987271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.987271Z digest=sha256:bdcc3adfde148a91a045e64123f6aaf0ddb353869f74d7bfe557c14f3b2e7d76

Observation 233e4718-6d55-4085-a644-6d508d92c1f0 · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Structured Pruning Learns Compact and Accurate Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.990020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.990020Z digest=sha256:80abb46334f2489f2659b7794ed6ff4086287560f170ebf3364c3d9ee21fa3b0

Observation 33c2b9fc-d032-40a7-95a7-0384e79a0ba4 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.992893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.992893Z digest=sha256:9d98886a1d9a5360b348a8f37962b925ecf03905a9ca987e4d1d082331322699

Observation 33c4be17-a40d-4aa6-935b-01e311898b3d · outbound

This paper cites ThinK: Thinner Key Cache by Query-Driven Pruning.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation ThinK: Thinner Key Cache by Query-Driven Pruning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:51.995190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:51.995190Z digest=sha256:c29065f72d37bd51854e84dfa934533a59e539cafa98f56858f10cd3fc5b8687

Observation 0273513d-35ee-4930-a305-18c2edda9540 · outbound

This paper cites X-pruner: explainable pruning for vision transformers.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation X-pruner: explainable pruning for vision transformers

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.424277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:51.998085Z digest=sha256:040a7314f3b4860ff9bf1fd52735135e30bb3a1d9a6dad99386adee15849505b

Observation 93c0355e-f848-45db-9b2f-12e6afd37318 · outbound

This paper cites Unified Visual Transformer Compression.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unified Visual Transformer Compression

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.000864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.000864Z digest=sha256:961cb40fc7efaf06f67404b2edf0c481a2ac9697cddd98298f352f5f462a382c

Observation 11a5dcf9-149d-4f7b-b713-3b8c2f8a1686 · outbound

This paper cites MoEfication: Transformer Feed-forward Layers are Mixtures of Experts.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEfication: Transformer Feed-forward Layers are Mixtures of Experts

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.004555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.004555Z digest=sha256:f3adc79d16360fc296903dcf1c7c6712edfea9bc78fea570b23832d6c768483b

Observation caea50ba-6c6e-424d-bb5c-0cc1ee2e5027 · outbound

This paper cites Knowledge distillation based on transformed teacher matching.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation based on transformed teacher matching

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.414808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:52.007595Z digest=sha256:d6e8946b9c57ada43a1c99c61b3c7cc0926db57f0c973bd9dedc26ad4f4be800

Observation 80958d0c-b3c7-4749-8de3-4432ec47e4ca · outbound

This paper cites The clock and the pizza: Two stories in mechanistic explanation of neural networks.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The clock and the pizza: Two stories in mechanistic explanation of neural networks

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.010319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.010319Z digest=sha256:94ed62c846e317fecde6f11001f573812e30b2b85236fb9f93f5ac65af737909

Observation 19db123c-225d-4678-83db-042e11c80ad4 · outbound

This paper cites LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.013139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.013139Z digest=sha256:6f050ea01cf78c1c9b7991bb0bc855e6df2e23e93b66ab9397dfcb0ea88ce881

Observation e89b4d73-6cb9-4692-bfff-0be0ab6974aa · outbound

This paper cites MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T14:38:52.016407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:38:52.016407Z digest=sha256:10855d0ee06650ab6459b14869128cabb155c703211847f4f22557425c4806a3

Observation 2a226cfb-df33-4745-a251-066a3616abc2 · outbound

This paper cites Possible solutions could include:.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Possible solutions could include:

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:38:52.402343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:52.019705Z digest=sha256:da8a77065f0b04d75f751bf4506316cecb75d7f9f5bc6fb80d314634d872cee9

Observation da8cde43-a337-4086-a708-a2e5e522a007 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.394336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:52.022253Z digest=sha256:dee570465e236779eb08f0c976d13de8318fb53e98eb51c95ad167cae4e76a8b

Observation 35c4f703-c447-4c33-ad9e-f9b61206eaa1 · outbound

This paper cites an unresolved cited work.

ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:38:52.385463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:38:52.024717Z digest=sha256:b0034caafc14ab0ec356fd9b74ccdc7dbd2dda5c1895d639a4430c262e1476ef

Pith citing papers

No inbound Pith citation observations are available.