Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:03.727373Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2509.10530.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:03.727373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b21f112d-6e84-4952-a97d-2b5314cff2ee · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 418fec02-e729-4cc7-b677-502c07e9dccf · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2cfd6db-3aca-4ed1-a360-c97afcc8c9aa · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Scaling Laws for Neural Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370dd306-3b37-448e-8b88-6fbb2d425ed9 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f772533-9eaf-4953-b708-019590d4c3b6 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c026a01c-5db7-416b-af46-ef1937360e92 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07fe1b71-d834-4af7-9164-0fc1304e57e5 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Adaptive mixtures of local experts
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b9c3a84-db0d-4c4a-a20f-374f2e2e96e0 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts What Does BERT Look At? An Analysis of BERT's Attention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2464f07-ac6f-458e-b872-bbbfc014d0b7 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 845cbdcc-ed51-4c1f-bd48-6711d3a0c56c · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Bert: Pre-training of deep bidirectional transform- ers for language understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 380e7d19-998e-4c41-86d2-dc1cbe62254f · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatgpt: Optimizing language models for dialogue, Nov 2022
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation bc83fb04-e1a6-4ac1-905c-270fe3c952d3 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts LLaMA: Open and Efficient Foundation Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d0b965-b4b4-4de5-b6ca-918a8a3aaf06 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c2cb9a-e169-4e83-92dc-b58f18f78bbf · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 722a6d10-20c9-480f-9b15-5bc66e284559 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcdbf71e-4747-4e51-ab4e-fb3743e3a3ea · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Mixture-of-Experts with Expert Choice Routing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86591840-0333-4746-b4ed-fd9cdb30bf34 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Palm: Scaling language modeling with pathways
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72e07bab-da7d-412b-82f1-329f27358bbf · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Efficiency Misnomer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f5f81b-586a-43af-821e-6494e872ac9f · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382d5947-2e0c-49bb-bc48-93a6cff44cf3 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297, 2020
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990c5cbe-0781-49b6-a3b7-c4fe2765a371 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Long Range Arena: A Benchmark for Efficient Transformers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7036773e-43d0-4468-a9c7-c562bbda26f5 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Base layers: Simplifying training of large, sparse models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b4d9c3d-c9c6-486e-8b0b-6a74e6c8c037 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Sparse is enough in scaling transformers.Advances in Neural Information Processing Systems, 34:9895–9907, 2021
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a66adfdd-9e1a-4cc7-ae92-830454ea166e · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hash layers for large sparse models.advances in neural information processing systems, 34:17555–17566, 2021
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62556ec8-19e2-4526-b831-7f93dfe37da5 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Synthesizer: Rethinking self-attention for transformer models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0554f798-9c13-4e51-881c-71471344b083 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Random Feature Attention
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da608b2d-cd50-4437-8ab0-3fccb3b16c01 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Skyformer: Remodel self-attention with gaussian kernel and nystr\" om method.Advances in Neural Information Processing Systems, 34:2122–2135, 2021
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f8d87c-1034-4894-a51a-7c7ba397f406 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3dd0fbf-0933-4c95-937e-b3e43f08fe74 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d273da-fc0a-4fcb-811e-48df01efc227 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Flex- moe: Scaling large-scale sparse pre-trained model training via dynamic device placement.Proceedings of the ACM on Management of Data, 1(1):1–19, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 89e83ad8-8caf-4f34-9518-5d356499d39d · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Go wider instead of deeper
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 14c12484-f515-47ff-be29-bd1b956b6f9f · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bdc8378-f738-437a-8556-561e68ffe5d0 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficientnet: Rethinking model scaling for convolutional neural networks
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62fac901-0581-4402-8615-6888bd3ccf09 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b8ed44-bf7b-4baa-891a-8711216a334a · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Dolan and Chris Brockett
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21cc405f-e474-4077-9ae3-3fe9aaad2f20 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficient Large Scale Language Modeling with Mixtures of Experts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b27b0d-4fc7-4af9-a25c-c2708bd6bcc1 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f302862-a9a5-4880-bb6a-98b23e3ba248 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linformer: Self-Attention with Linear Complexity
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e832ac-c52b-4e3c-be74-d4fbef186d24 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Reformer: The Efficient Transformer
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 923ac743-556d-46f9-a7ea-c4b6c7c65b90 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Rethinking Attention with Performers
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1f67ac-16a8-476c-82fa-13aa06552c53 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b74526-d5c2-4629-bf7c-b2179ccb8030 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fea4b75b-7025-4fd8-b160-97d4ce5e80af · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67704a54-058a-4ec7-8875-311ae2f4964d · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Superglue: A stickier benchmark for general-purpose language understanding systems.Advances in neural information processing systems, 32, 2019
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2e7070-68de-4329-a7f9-3679667f5e83 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations, 2021
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7767dbbf-f17c-4873-bfe0-d3410182ba68 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts CMMLU: Measuring massive multitask language understanding in Chinese
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ccf242e-1ea1-4eeb-9979-d6badaa23c22 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20cff156-4cbb-4f87-9db4-a7e43b3891fa · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef53fdfa-93b6-4e3a-9271-161471d16a54 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Training Verifiers to Solve Math Word Problems
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc47581-586c-4aff-8f88-cc0b5bc31742 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring Mathematical Problem Solving With the MATH Dataset
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b3adbe-ecf5-4878-91fb-9c0b0dc9c85a · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Program Synthesis with Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5338c63-003a-49a5-83da-7e989e0e3965 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Evaluating Large Language Models Trained on Code
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d36177ae-036b-41b6-8caa-36cf8829469b · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Qwen3 technical report, 2025
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23e9b094-a44d-403a-a3ec-2ea2b7caa112 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aec8be8-f622-4277-958d-21d921fe120e · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Gemma 3 technical report, 2025
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 5bcd0223-bb1a-419a-9456-10ab6af55da1 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Llama 3 Herd of Models, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 37a12c06-9543-4e0c-bb9f-ddcf54f92780 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hewett, Mojan Javaheripi, Piero Kauffmann, James R
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f61c776-484d-4836-ab5f-6d27171eb856 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Multi-Scale Dense Networks for Resource Efficient Image Classification
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 368fc9a0-32e7-48e7-81f4-69853d82ec85 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Manning, Andrew Ng, and Christopher Potts
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4450c4-38db-4f31-b009-ecea4e70e44d · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6e7cf706-c3cc-450b-9e5a-57ffaaadb444 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a72df6b2-87ca-4b4b-a0f8-eb6ad8848330 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Squad: 100,000+ questions for machine com- prehension of text
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f607e9a1-5cd4-4e2f-84d4-3f398bb577c9 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The pascal recognising textual entailment challenge
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f91b630-358f-4061-8801-cb20e554f439 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts PaLM 2 Technical Report
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9d747b-9a42-4583-9d1b-02b07b4e6363 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linear attention is (maybe) all you need (to understand transformer optimization)
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc30f7b7-3e5a-4e4a-80ed-637a98deb910 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440c65a9-df59-4eaa-99f9-ad67e76f2045 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts FNet: Mixing Tokens with Fourier Transforms
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 598c47b0-3365-4852-a613-acafa17a55f4 · outbound
Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Densely Connected Convolutional Networks
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.