Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:00.523197Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2507.08771.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:20:00.523197Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T22:50:51.900169Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:39:56.501546Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25fa0fb7-6e23-4bfd-8cc3-f9f36670b563 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PIQA : Reasoning about physical commonsense in natural language
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b57dedeb-ae76-41c5-b3a2-f55d69df0f44 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d59b9e6c-aee0-49b5-b253-a7558066ea20 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity BoolQ : Exploring the surprising difficulty of natural yes/no questions
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d36ce6a-fa90-48ad-8658-51e69906c554 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity TyDi QA : A benchmark for information-seeking question answering in typologically diverse languages
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fc253894-037c-4f0e-9518-3778fdd8336b · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b66a0274-5a9c-40fd-a5c4-936255cadfbd · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Language modeling with gated convolutional networks
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation becf4859-0d6b-47cb-a1c7-33fad50d7330 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 086fbc5d-070c-4a7f-97f1-51f4fe6b4ca0 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity The Llama 3 Herd of Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20c700f7-2af5-4d32-9a65-1da0ef4ef76f · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Switch Transformers : Scaling to trillion parameter models with simple and efficient sparsity
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8dd14c3c-02c3-43d8-a52c-63e68d5035e8 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity SparseGPT : Massive language models can be accurately pruned in one-shot
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9883f86a-5a9f-4832-bfc4-db0bda45f159 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity MegaBlocks : Efficient sparse training with mixture-of-experts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 709b6dab-f0d3-4921-a1af-d33a2f5ba256 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e61f35ce-cdea-4b3b-83f8-1e1b304a7f98 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity MiniLLM: On-Policy Distillation of Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f82d564-9f3c-467a-876f-ff097377227b · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity FastMoE: A Fast Mixture-of-Expert Training System
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4885e582-1654-482f-ba39-538194b9bca2 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c560bb00-3ef7-4756-a726-d4c3e909313b · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4a7d469-e966-483b-81f5-98417631cdf9 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Harder tasks need more experts: Dynamic routing in MoE models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fe1ec6-c0e1-40e4-9fae-4a7ee2be6c00 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Tutel: Adaptive mixture-of-experts at scale
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1b3dc5d7-1be0-492d-8bb4-a90ea8d92b04 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Mixtral of Experts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7060e794-53d6-4614-9403-522dba46b907 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Scaling Laws for Fine-Grained Mixture of Experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e6cce0-03e7-4d88-9e45-03f35124285f · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Fast inference from Transformers via speculative decoding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a0488d49-b902-4415-a79d-6c3bb7b0083d · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity StarCoder: may the source be with you!
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 516aafb3-7494-4d74-b9cd-4b1998290d42 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c78072f2-a450-4ca3-9f8a-8033e5a7460d · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity EAGLE-2 : Faster inference of language models with dynamic draft trees
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d03b74b2-0cf4-4847-bd92-31c8cb1afab9 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity The lazy neuron phenomenon: On emergence of activation sparsity in Transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f4f91eca-3bb6-465d-aeff-540963d4cddc · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e398f203-85dc-44a6-81fa-8cc2a9e7ff9a · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity DeepSeek-V3 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47be00ea-9acd-4eb0-a587-56d76c9d458e · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity GRIN: GRadient-INformed MoE
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7468f16-caaa-4d13-a243-fb27d8bae376 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Deja Vu : Contextual sparsity for efficient LLMs at inference time
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 72eb8b80-925b-4b3b-8508-d59bd534c81f · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Sparsing Law: Towards Large Language Models with Greater Activation Sparsity
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3d2b29-e9fa-4350-8b6a-64625e583a2e · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity LLM-Pruner: On the Structural Pruning of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dd08c2-2153-43ea-a3d4-cf8960f2ad0a · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c792a5f-a297-4f88-a602-aff13ae896e4 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Soft Merging of Experts with Adaptive Routing
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 849baa19-38b4-4075-bc2a-937cad948e88 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 575f85a8-5179-4854-a603-839825c544cb · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Exploring the limits of transfer learning with a unified text-to-text Transformer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fd5e66af-7322-43de-a68a-59eb8cc327b5 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Improving Dictionary Learning with Gated Sparse Autoencoders
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfb9ccb-c0d2-412c-955c-cf9a429fa4c8 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Searching for Activation Functions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ad04899-f78c-420b-afd3-cc161c9d8e16 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity SocialIQA : Commonsense reasoning about social interactions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9d3363eb-a434-4907-826e-0a48c4902190 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a608e87e-f83b-4a60-8507-4b651100694d · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity GLU Variants Improve Transformer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f5ddc7-99d1-45d0-8b5b-cd3b6b779129 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b190fc45-d9e5-4ddf-9ad9-dac4c51e70e5 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity P ro S parse: Introducing and enhancing intrinsic activation sparsity within large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0aeb7da2-92c1-4501-bf32-4b3b6ad58c68 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e03d1162-e296-44da-a76d-51bfbd2de612 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Turbo Sparse: Achieving LLM SOTA Performance with Minimal Activated Parameters
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2649523e-6833-4eb3-ae6b-9b1de0c4042f · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity A Simple and Effective Pruning Approach for Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365695a0-91c8-4964-a070-4aa6c8fbb17a · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity CUTLASS , Jan 2023
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aaabc245-9903-425a-a86a-2377bfe0ce71 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Efficient large language models: A survey
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e66add9-f76c-47ba-ab8f-76dc9a53554e · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad462045-dcd7-40d1-8b27-cfd91c2ebd0c · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ReMoE: Fully Differentiable Mixture-of-Experts with ReLU Routing
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cece94a5-82ba-4719-b580-25b15f3b141d · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Magicoder: Empowering Code Generation with OSS-Instruct
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4934800a-3021-4a03-83c0-93cf3af0f586 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3a1fcbb4-978f-4882-940b-574372928e36 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4fe0660-d3de-4f05-8bcb-269058477b56 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9b26c98d-7cb2-413b-9054-b3e19303c0fa · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bd469b-6ec3-40fb-a33e-85f97cc595f6 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity PowerInfer-2: Fast Large Language Model Inference on a Smartphone
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a97201bc-28ae-4c43-ab1a-f42024abd097 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25797477-e26a-421e-9ef0-e44c03e8bedc · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ZeroQuant-V2: Exploring Post-training Quantization in LLMs from Comprehensive Study to Low Rank Compensation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a277af18-f8a9-4d11-942e-f0bcdd3958cc · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity HellaSwag : Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pp.\ 4791--4800, 2019
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c5e6c3cf-3197-4da6-b48a-e310b78750f1 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Root mean square layer normalization
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 01cbe2a3-37ea-4fb4-8e35-e1dc8e9e3ef6 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9db0349-ea9d-4d1e-b7d4-e71de08952b2 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Exploring the benefit of activation sparsity in pre-training
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 516e2264-3414-4840-8cbc-e056a509ae59 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Ouroboros: Generating longer drafts phrase by phrase for faster speculative decoding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3d8819fc-ab7a-49f8-aae5-4dc2a5f00633 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative Sampling
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec30a20a-b760-4d15-8bf7-5a8616bae220 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11de79ca-38ea-4612-a47d-3447bd797cfc · outbound
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a79635e-532f-4ff6-88bd-ab19558b5e33 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Unresolved cited work
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefe486b-021e-4a95-a1ab-6965acf4db97 · outbound
BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity Unresolved cited work
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443bc417-9566-4821-9188-faa0047dcff7 · inbound
dMoE: dLLMs with Learnable Block Experts BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5527f436-88e0-4a7d-bd8e-e1decd92f1cc · inbound
GeMoE: Gating Entropy is All You Need for Uncertainty-aware Adaptive Routing in MoE-based Large Vision-Language Models BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.