Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2411.15281.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T14:38:52.024717Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f3d4b98e-6203-4ea3-84f8-481944c31916 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e161dd6-471d-400d-a50f-9a5dba66d90e · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Fluctuation-based adaptive structured pruning for large language models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02aab24a-4e03-4b51-b3af-43518215b7cd · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamic context pruning for efficient and interpretable autoregressive transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49bea46-440c-4585-bf77-5a4d907f760d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A general language assistant as a laboratory for alignment, 2021
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4aa84aa-226c-415f-94ac-a90b796e9a9c · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mitigat- ing open-vocabulary caption hallucinations, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac38b26-ef54-4f65-9184-f54dbebf4816 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensorflow, March 2021
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfd730a8-8e37-4c5e-9891-b058b4c1b3cf · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Token Merging: Your ViT But Faster
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b91a4d79-f0b6-4bf1-a0bb-ffe534159e3a · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Emerging properties in self-supervised vision transformers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92479429-200b-4edf-9fc0-11b80eed0654 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Vision transformer slimming: Multi-dimension searching in continuous optimization space
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b13b73c1-4e1b-42e8-b792-8ebfafee7847 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01cc1d15-c44a-4cdf-a841-8966e90133a5 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The lottery ticket hypothesis for pre-trained bert networks
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f86d788e-745b-4e07-9c79-257d048ac7ca · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The principle of diversity: Training stronger vision transformers calls for reducing all levels of redundancy
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fee15ab5-e6a7-492a-9566-7a05fb5d132d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A toy model of universality: Reverse engineering how networks learn group operations
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd75c708-4a6f-4750-a7f1-62719a01e8af · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training Verifiers to Solve Math Word Problems
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3996ee05-cad3-4b1d-bff1-23af3959df71 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbb301ab-1bab-4673-8ae0-54f5dd6d3740 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Analyzing Redundancy in Pretrained Transformer Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5538669b-0403-47af-876a-931014f27133 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Imagenet: A large- scale hierarchical image database
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 650b0c88-dbfe-41b6-b3cb-405e0417c8d8 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Qlora: Efficient finetuning of quantized llms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9609917-7f82-4be9-a255-a1b9eef0f188 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Eventful transformers: leveraging temporal redundancy in vision transformers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a1f76e72-8720-422c-a01e-9072164333cf · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A mathematical framework for transformer circuits
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199df78c-ee4c-4a08-bc03-183021e11d32 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Depgraph: Towards any structural pruning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4907c096-c9f4-4529-82a1-5959b9676cdd · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity, 2022
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1c5090-53ce-408f-834f-467a84021cfa · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Transformer Feed-Forward Layers Are Key-Value Memories
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ba1185-ab48-45e3-9583-0794cbcd0c7e · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Successor Heads: Recurring, Interpretable Attention Heads In The Wild
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26169097-f8ff-4db5-af4b-4befc0b14ab3 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MiniLLM: Knowledge distillation of large language models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 414e3cc2-5e12-4c2a-9203-cec92ff814e8 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Learning efficient vision transformers via fine-grained manifold distillation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 131feb79-a177-4bfb-a74d-01897e54cf75 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Masked autoencoders are scalable vision learners, 2021
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7485a9dc-058c-473a-8750-49079323c695 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation What Matters in Transformers? Not All Attention is Needed
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb07f159-3abb-4775-ac42-e7edb1f28aaf · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Distilling the Knowledge in a Neural Network
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 390edfd7-ede8-4410-9e30-cc7a3240712d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sparse Progressive Distillation: Resolving Overfitting under Pretrain-and-Finetune Paradigm
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ad95b5a0-f5e3-493d-a44e-54b6f12f8fb1 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture of Nested Experts: Adaptive Processing of Visual Tokens
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3be0e74-894d-484d-9fa7-405a2e6d81d2 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distillation into self-attention heads for improving transformer-based end-to-end neural speaker diarization
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55eab72c-802c-498f-82b7-411ff9c34e9a · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Expedited training of visual conditioned language generation via redundancy reduction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 084e9234-9b0f-4e80-addb-0417bb2a341c · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixtral of Experts
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14383e2d-b807-48c9-b653-95ee6e53d585 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-supervised 3d anatomy segmentation using self-distilled masked image transformer (smit)
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 682ee29e-5d31-4b4a-9497-a16b0031564e · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation TinyBERT: Distilling BERT for Natural Language Understanding
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74b67fb5-c03e-424c-8183-147344f8fe73 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e73ced7e-ee60-4cf1-998e-531e43d0c297 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-Distillation for Further Pre-training of Transformers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4113c3f4-3216-4e76-a6e1-48836cd0a659 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Clustered imagenet labels for training production-friendly image classifier
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5593dda4-a275-4e98-a5f3-4d9a0ef3351a · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation via the target-aware transformer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b84504f-67ac-4341-814b-58b247eef34d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50ac9757-c5a3-4dfb-8f42-4646693131ec · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Visual instruction tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aa03e23-d79f-447a-8594-7510f5fce33e · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Oscillation-free quantization for low-bit vision transformers
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a0cf4fec-e8a2-451e-b519-75c6bbb1ef38 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Post-training quantization for vision transformer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106afdf3-ea68-44a8-a341-4635adc29ec7 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Anytime Dense Prediction with Confidence Adaptivity
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7688bba1-7d21-40ad-8772-053af3629bdb · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A transformer-based model with self-distillation for multimodal emotion recognition in conversations
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 89d96134-36af-4bb5-a78c-36ab88e259ad · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Llm-pruner: On the structural pruning of large language models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc78494-02f8-4e58-ba88-8ed448c12c59 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Copy Suppression: Comprehensively Understanding an Attention Head
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bd13a09-5670-464f-a3aa-340dc8418f03 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Locating and editing factual associations in gpt
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa4206b-3f0a-453a-ae07-76d0bcfe2ca4 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Circuit Component Reuse Across Tasks in Transformer Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2baf1a4b-320a-45e0-82ef-de20738605bb · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Zoom in: An introduction to circuits
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad33b7e-34c7-416e-887b-3928d987a573 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation In-context Learning and Induction Heads
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f652ae6a-d2fb-4fc5-abb2-6ba6a042452c · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Ia-red 2: Interpretability-aware redundancy reduction for vision transformers
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5add9e11-2b4b-45b8-be8d-75586f889ba1 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-evolving vision transformer for chest x-ray diagnosis through knowledge distillation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a20e40e5-ec2e-483d-b558-a50e0fc0c57f · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A practical review of mech- anistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c018973-c0ee-4a62-9d21-bf2519e8e086 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Dynamicvit: Efficient vision transformers with dynamic token sparsification
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f16e29e-3e9a-41c1-9fcd-3cabd2f40520 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf3dd24-01e7-4af5-9c2b-c8b570da2546 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39d689fd-4cd5-4e07-8e20-31bc538d89af · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Confident adaptive language modeling
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 391d1b87-f156-45bc-b39e-55f78eb52bd6 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e457abb-0556-4269-abc2-bfc9f1873c4e · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e6f198-ea37-4b4c-954b-1d271ad9dd34 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 697491b3-5265-46f7-88e0-6c8627f8810b · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tasked: transformer-based adversarial learning for human activity recognition using wearable sensors via self-knowledge distillation
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73e114a5-6a81-4076-92a5-f43ba06e4e2d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Self-distilled vision transformer for domain generalization
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d3b10aa2-14bb-43f3-b59c-0b57d06a790a · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patient Knowledge Distillation for BERT Model Compression
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8312c520-0e96-41d9-b045-14fd40262c2b · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b47718a4-6be5-4669-90a5-67c84787c257 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Patch slimming for efficient vision transformers
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc3cce1a-9e23-42a2-8fba-d94c5de0578b · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8378a85f-e41e-4276-9c62-43ec2ac1af75 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Training data-efficient image transformers & distillation through attention
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3fa46d1-7012-4a75-910a-88ec06df7c40 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Attention is all you need
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03c70042-9d43-43c0-8c76-c5f78c24529f · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Rkld: Reverse kl-divergence- based knowledge distillation for unlearning personal information in large language models, 2024
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8adb33bb-3086-440e-b335-7cab20402372 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62e5c21-9f43-4b81-ab10-ab17702fb178 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e24b8f23-3957-4d42-9397-6e0ec64b77d3 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Last: Label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1749f87a-2a79-4e35-ba69-ae658f8d1b8d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Tinyvit: Fast pretraining distillation for small vision transformers
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 23cc25b4-72e7-4e58-bf94-be7ee936f505 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233e4718-6d55-4085-a644-6d508d92c1f0 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Structured Pruning Learns Compact and Accurate Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c2b9fc-d032-40a7-95a7-0384e79a0ba4 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33c4be17-a40d-4aa6-935b-01e311898b3d · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation ThinK: Thinner Key Cache by Query-Driven Pruning
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0273513d-35ee-4930-a305-18c2edda9540 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation X-pruner: explainable pruning for vision transformers
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93c0355e-f848-45db-9b2f-12e6afd37318 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unified Visual Transformer Compression
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11a5dcf9-149d-4f7b-b713-3b8c2f8a1686 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEfication: Transformer Feed-forward Layers are Mixtures of Experts
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caea50ba-6c6e-424d-bb5c-0cc1ee2e5027 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Knowledge distillation based on transformed teacher matching
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 80958d0c-b3c7-4749-8de3-4432ec47e4ca · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation The clock and the pizza: Two stories in mechanistic explanation of neural networks
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19db123c-225d-4678-83db-042e11c80ad4 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89b4d73-6cb9-4692-bfff-0be0ab6974aa · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation MoEBERT: from BERT to Mixture-of-Experts via Importance-Guided Adaptation
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a226cfb-df33-4745-a251-066a3616abc2 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Possible solutions could include:
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da8cde43-a337-4086-a708-a2e5e522a007 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35c4f703-c447-4c33-ad9e-f9b61206eaa1 · outbound
ElastiFormer: Learned Redundancy Reduction in Transformer via Self-Distillation Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
No inbound Pith citation observations are available.