Pith. sign in

Paper Citation Record · LEDGER

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts

As of 16 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2509.10530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10530 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:32:03.727373Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b21f112d-6e84-4952-a97d-2b5314cff2ee · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.102282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.102282Z digest=sha256:0af3dcbd234ae3cd8ce83a867cb1857a43161afbc6c57c3c1286edea98b1dabb

Observation 418fec02-e729-4cc7-b677-502c07e9dccf · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.114581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.114581Z digest=sha256:1668033d85cbad6597df6865ee6c9c03bb074df840b2548837b23f7632bcb119

Observation a2cfd6db-3aca-4ed1-a360-c97afcc8c9aa · outbound

This paper cites Scaling Laws for Neural Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Scaling Laws for Neural Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.122984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.122984Z digest=sha256:90c287a307317ff40a8f456e0a7ea961209bfa39ddfba227b66ee8799738235f

Observation 370dd306-3b37-448e-8b88-6fbb2d425ed9 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.132827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.132827Z digest=sha256:1699c2df72da7cfed7c132b18bc11a2604f982b0666d3bacabd19712c5920542

Observation 7f772533-9eaf-4953-b708-019590d4c3b6 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.141013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.141013Z digest=sha256:4c217ca240a3d054ccc82b365791960a185019a206147bd15384facdafcf3d53

Observation c026a01c-5db7-416b-af46-ef1937360e92 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.152543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.152543Z digest=sha256:0e77bf0743546c4e53f30f5cb42d606edfc80839db8b4aa4c32c0ca55102250a

Observation 07fe1b71-d834-4af7-9164-0fc1304e57e5 · outbound

This paper cites Adaptive mixtures of local experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Adaptive mixtures of local experts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.161637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.161637Z digest=sha256:91eabba5c788248cfbe5e0a460747cb2c25609caef533fe4624169013da1650c

Observation 0b9c3a84-db0d-4c4a-a20f-374f2e2e96e0 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts What Does BERT Look At? An Analysis of BERT's Attention

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.168149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.168149Z digest=sha256:3cafefaaaba5f5f075c2463496e79b3884b4fbd5b029abc9cdb2072c5aa887c8

Observation e2464f07-ac6f-458e-b872-bbbfc014d0b7 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.174237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.174237Z digest=sha256:8be8def517af8efcf8d07fdaacde03fd9ed69d00e2fb8cba120d768c38baea51

Observation 845cbdcc-ed51-4c1f-bd48-6711d3a0c56c · outbound

This paper cites Bert: Pre-training of deep bidirectional transform- ers for language understanding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Bert: Pre-training of deep bidirectional transform- ers for language understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.178360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.181778Z digest=sha256:b5bbcaf243bbdbf3f60fe45b43a26754c168cf503585aa0907c23d931456c120

Observation 380e7d19-998e-4c41-86d2-dc1cbe62254f · outbound

This paper cites Chatgpt: Optimizing language models for dialogue, Nov 2022.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatgpt: Optimizing language models for dialogue, Nov 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.160707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.188074Z digest=sha256:e970def185ea2d8b0467f12acaa5ef6b1a0fc88fb9c70d88857cdd301521569e

Observation bc83fb04-e1a6-4ac1-905c-270fe3c952d3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts LLaMA: Open and Efficient Foundation Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.196590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.196590Z digest=sha256:ce65ad4c339e2746eab62cb9a43c6193bb5f2b3e72bf24af3556238ccffefca6

Observation e0d0b965-b4b4-4de5-b6ca-918a8a3aaf06 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.205921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.205921Z digest=sha256:25570ce14ca054a88ddbebf6ba61e20c9e1c893cccb10e9bc05504b9024ffd37

Observation 81c2cb9a-e169-4e83-92dc-b58f18f78bbf · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Glam: Efficient scaling of language models with mixture-of-experts

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.124029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.215508Z digest=sha256:1b09a976a94cb0332cf6f36a0f47e99cb91b47fcfb8e9c3f1884ca93b78e02a8

Observation 722a6d10-20c9-480f-9b15-5bc66e284559 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.226383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.226383Z digest=sha256:7865db6896ca51dffd851405b594f1ada2293f0d9f390c90ac7bde02c2a5a9c8

Observation bcdbf71e-4747-4e51-ab4e-fb3743e3a3ea · outbound

This paper cites Mixture-of-Experts with Expert Choice Routing.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Mixture-of-Experts with Expert Choice Routing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.233617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.233617Z digest=sha256:a737c263099fbc630629ccb6cd71e0adc7002721132aac2f8bfa624e45d28e42

Observation 86591840-0333-4746-b4ed-fd9cdb30bf34 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Palm: Scaling language modeling with pathways

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.249006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.249006Z digest=sha256:abb9034fc71337c3ad9fb5c97a8321daebe47c1b3e01b420f165d1813bad5b48

Observation 72e07bab-da7d-412b-82f1-329f27358bbf · outbound

This paper cites The Efficiency Misnomer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Efficiency Misnomer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.259828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.259828Z digest=sha256:43a7e75f1f195813d65beb9684ca1d4614061c0fd2ac87f7dffb44ac94018f9f

Observation b0f5f81b-586a-43af-821e-6494e872ac9f · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.268556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.268556Z digest=sha256:bafe255b208482957e6e51553b77c30536b13cac77d03fb95ca79dc0b84738da

Observation 382d5947-2e0c-49bb-bc48-93a6cff44cf3 · outbound

This paper cites Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Big bird: Transformers for longer sequences.Advances in neural information processing systems, 33:17283–17297, 2020

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.278164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.278164Z digest=sha256:393e9b8f4207ca11e0c027b222aa8c6f16df0358812cca54f947e5ecf3583c5f

Observation 990c5cbe-0781-49b6-a3b7-c4fe2765a371 · outbound

This paper cites Long Range Arena: A Benchmark for Efficient Transformers.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Long Range Arena: A Benchmark for Efficient Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.286629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.286629Z digest=sha256:d69c840297680defb493a754004c00bd40d590e3a71087ba7b100776e2c2d527

Observation 7036773e-43d0-4468-a9c7-c562bbda26f5 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Base layers: Simplifying training of large, sparse models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.070594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.293221Z digest=sha256:f949a026584721d0be774c4d5545df5505bf8416e32b5354101a75243af1a54b

Observation 9b4d9c3d-c9c6-486e-8b0b-6a74e6c8c037 · outbound

This paper cites Sparse is enough in scaling transformers.Advances in Neural Information Processing Systems, 34:9895–9907, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Sparse is enough in scaling transformers.Advances in Neural Information Processing Systems, 34:9895–9907, 2021

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:05.041059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.300884Z digest=sha256:f5d4a65ec57ab9d6b2a52f018c554be720d6a1e29069469a6fd805a8b70758f4

Observation a66adfdd-9e1a-4cc7-ae92-830454ea166e · outbound

This paper cites Hash layers for large sparse models.advances in neural information processing systems, 34:17555–17566, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hash layers for large sparse models.advances in neural information processing systems, 34:17555–17566, 2021

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.313419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.313419Z digest=sha256:8db9f2f7213534e74111e08ff538f8833437715f50216951704cc4e96cedcb7c

Observation 62556ec8-19e2-4526-b831-7f93dfe37da5 · outbound

This paper cites Synthesizer: Rethinking self-attention for transformer models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Synthesizer: Rethinking self-attention for transformer models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.994516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.323430Z digest=sha256:cabd0bab502e3f3f6d4a2c9995180c6f6cf2316d723dcd1d3688cd700da1aa16

Observation 0554f798-9c13-4e51-881c-71471344b083 · outbound

This paper cites Random Feature Attention.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Random Feature Attention

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.332917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.332917Z digest=sha256:6086cfbd20dcad333c06b04b571ee91cfb67df08fb9a5c3f86f3c9a2f80718e6

Observation da608b2d-cd50-4437-8ab0-3fccb3b16c01 · outbound

This paper cites Skyformer: Remodel self-attention with gaussian kernel and nystr\" om method.Advances in Neural Information Processing Systems, 34:2122–2135, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Skyformer: Remodel self-attention with gaussian kernel and nystr\" om method.Advances in Neural Information Processing Systems, 34:2122–2135, 2021

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.340929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.340929Z digest=sha256:94683ab0d5064f795d8fc26458dabed1fcd882e35ba600681950740d9db4fff3

Observation 58f8d87c-1034-4894-a51a-7c7ba397f406 · outbound

This paper cites Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Fastermoe: modeling and optimizing training of large-scale dynamic pre-trained models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.354871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.354871Z digest=sha256:1f3b0d1f6a4cd3f46c7ed4d1e4da7fd173c07b852dfb6d92ccfb8a1baae2954c

Observation d3dd0fbf-0933-4c95-937e-b3e43f08fe74 · outbound

This paper cites MegaBlocks: Efficient Sparse Training with Mixture-of-Experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MegaBlocks: Efficient Sparse Training with Mixture-of-Experts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.365892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.365892Z digest=sha256:ae69d596ecf58115d6cf994a89c06be730344433a2aacec5b09e01150886a9a1

Observation e0d273da-fc0a-4fcb-811e-48df01efc227 · outbound

This paper cites Flex- moe: Scaling large-scale sparse pre-trained model training via dynamic device placement.Proceedings of the ACM on Management of Data, 1(1):1–19, 2023.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Flex- moe: Scaling large-scale sparse pre-trained model training via dynamic device placement.Proceedings of the ACM on Management of Data, 1(1):1–19, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.947004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.381550Z digest=sha256:228d76290e8d0d3ada453ebc921662a9193464d95892905aa31780973ed43b84

Observation 89e83ad8-8caf-4f34-9518-5d356499d39d · outbound

This paper cites Go wider instead of deeper.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Go wider instead of deeper

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.923186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.388315Z digest=sha256:5f5e762ef2f5d49cc1052783fbc94b8c133f44c275a78c9cd0af5e73a48b7c57

Observation 14c12484-f515-47ff-be29-bd1b956b6f9f · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.398538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.398538Z digest=sha256:2311f611b00e666c0b1181435d49ba057e5862a2105a27ed056383624f4d8fa2

Observation 9bdc8378-f738-437a-8556-561e68ffe5d0 · outbound

This paper cites Efficientnet: Rethinking model scaling for convolutional neural networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficientnet: Rethinking model scaling for convolutional neural networks

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.406558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.406558Z digest=sha256:83505592c2d00ac78e7356eddf3118a34b507ad29e618dda637f2399ea0d8e98

Observation 62fac901-0581-4402-8615-6888bd3ccf09 · outbound

This paper cites MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.413681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.413681Z digest=sha256:d9ab024d8b7613b0023d9f22d71c1e05f5c6ddb146cc519f1f06fd51213f3169

Observation 22b8ed44-bf7b-4baa-891a-8711216a334a · outbound

This paper cites Dolan and Chris Brockett.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Dolan and Chris Brockett

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.422607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.422607Z digest=sha256:5ff9a7d12faa60265691a20724cd47076f5a242c58d310369542f9e02feb13ae

Observation 21cc405f-e474-4077-9ae3-3fe9aaad2f20 · outbound

This paper cites Efficient Large Scale Language Modeling with Mixtures of Experts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.432673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.432673Z digest=sha256:fe653c29bee25be6edaaacf9fdf5fca77c87c9c7baf047abba2565ff9380c7a1

Observation d8b27b0d-4fc7-4af9-a25c-c2708bd6bcc1 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.440111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.440111Z digest=sha256:11d0b4b295d703fce0d1e58149d6c2e9c42eb8e253b858728b711374ca59616a

Observation 6f302862-a9a5-4880-bb6a-98b23e3ba248 · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linformer: Self-Attention with Linear Complexity

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.462100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.462100Z digest=sha256:90d1d5684f592619d29934089e7b3081d461b110e6472fb22e1b0b312e3610a0

Observation 94e832ac-c52b-4e3c-be74-d4fbef186d24 · outbound

This paper cites Reformer: The Efficient Transformer.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Reformer: The Efficient Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.473597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.473597Z digest=sha256:765a4d0a5b912334dcf42aa602f6d117f821b03ce1c0ac0c3d58c59c92414e87

Observation 923ac743-556d-46f9-a7ea-c4b6c7c65b90 · outbound

This paper cites Rethinking Attention with Performers.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Rethinking Attention with Performers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.482557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.482557Z digest=sha256:244424c7ebba634f66a246c825f317794a96bacd68565738c154c506c69ef28e

Observation 9c1f67ac-16a8-476c-82fa-13aa06552c53 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.505998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.505998Z digest=sha256:33e1987d0b5b733be00568f409e2dc269fe0d03857746aab45b876e40bfa968b

Observation 39b74526-d5c2-4629-bf7c-b2179ccb8030 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.516804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.516804Z digest=sha256:06a4263e1d936a0c9b304d84efb1a4d63d7a0a1a8aa8bdb4e650236ed9d96ff0

Observation fea4b75b-7025-4fd8-b160-97d4ce5e80af · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.530148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.530148Z digest=sha256:9232d935d2f5b896e0833158c287b4d8a5ec164dc7cadbc3da475a959c1692d6

Observation 67704a54-058a-4ec7-8875-311ae2f4964d · outbound

This paper cites Superglue: A stickier benchmark for general-purpose language understanding systems.Advances in neural information processing systems, 32, 2019.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Superglue: A stickier benchmark for general-purpose language understanding systems.Advances in neural information processing systems, 32, 2019

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.542392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.542392Z digest=sha256:32e6ad7d106228340b9ccf07a4b7ad8c2680a09d23c2e091e63bb6adf064abd8

Observation eb2e7070-68de-4329-a7f9-3679667f5e83 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations, 2021.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations, 2021

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.842222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.549030Z digest=sha256:f8c7aa36e3b6cb743934a9352ca56fc1838891479dcfda6c467a9e1d64482a49

Observation 7767dbbf-f17c-4873-bfe0-d3410182ba68 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts CMMLU: Measuring massive multitask language understanding in Chinese

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.557780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.557780Z digest=sha256:f930bbed326e2f4f2e16a473c78e56a045e5a96763d5d55b4a22da594d35d67c

Observation 4ccf242e-1ea1-4eeb-9979-d6badaa23c22 · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.566309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.566309Z digest=sha256:f86c1727044935cbb9abb4ad8fe4f42f8a5cb317f7435448a6e33481c911ec6a

Observation 20cff156-4cbb-4f87-9db4-a7e43b3891fa · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.573573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.573573Z digest=sha256:7e6d4c8617b57aba0e3cc53ee9583d55d67679f28a2e1f3f7345d21d4f96a54b

Observation ef53fdfa-93b6-4e3a-9271-161471d16a54 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.580103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.580103Z digest=sha256:1442a73449920f67730b0eebd31d859aa9347a7727282aa49ff2514d95978c91

Observation ebc47581-586c-4aff-8f88-cc0b5bc31742 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Measuring Mathematical Problem Solving With the MATH Dataset

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.587345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.587345Z digest=sha256:790ac1050306f0ce889a37113e85e79879e0c077af3970187a7347e58f8939c0

Observation 81b3adbe-ecf5-4878-91fb-9c0b0dc9c85a · outbound

This paper cites Program Synthesis with Large Language Models.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Program Synthesis with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.598148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.598148Z digest=sha256:6f1f0958e6a14c708b8140bd6668b96efd17a19fa8c0ea0f85c2bfef7faeec9e

Observation f5338c63-003a-49a5-83da-7e989e0e3965 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Evaluating Large Language Models Trained on Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.605123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.605123Z digest=sha256:d0c898e3d452ad8f6bc5b935b25f71824dae7aa457bfa0f7f0b0349fd99c5dc8

Observation d36177ae-036b-41b6-8caa-36cf8829469b · outbound

This paper cites Qwen3 technical report, 2025.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Qwen3 technical report, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.613500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.613500Z digest=sha256:18886ab0319dd9e5081b6ccc906976866a06921cf22af03ce8d0e54d95c5d892

Observation 23e9b094-a44d-403a-a3ec-2ea2b7caa112 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.619572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.619572Z digest=sha256:4bab379f67d68b657a9c95ef3416fe613df526ce0397384ddd5d11edfd609bee

Observation 4aec8be8-f622-4277-958d-21d921fe120e · outbound

This paper cites Gemma 3 technical report, 2025.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Gemma 3 technical report, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.798914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.625468Z digest=sha256:0e969561df845113454010d4370d53720be8b1c45aa8acadbd645650397cc942

Observation 5bcd0223-bb1a-419a-9456-10ab6af55da1 · outbound

This paper cites The Llama 3 Herd of Models, 2024.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The Llama 3 Herd of Models, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.777934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.631773Z digest=sha256:e77334442414b656adcf1f92498d4a97287065e66e2f160a26c40af81c8fd623

Observation 37a12c06-9543-4e0c-bb9f-ddcf54f92780 · outbound

This paper cites Hewett, Mojan Javaheripi, Piero Kauffmann, James R.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Hewett, Mojan Javaheripi, Piero Kauffmann, James R

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.754191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.640184Z digest=sha256:4c667e9f93782360a15a6d79fe62878fb86e5b38cf8733cb8c00163a8518cab9

Observation 8f61c776-484d-4836-ab5f-6d27171eb856 · outbound

This paper cites Multi-Scale Dense Networks for Resource Efficient Image Classification.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Multi-Scale Dense Networks for Resource Efficient Image Classification

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.650932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.650932Z digest=sha256:e808ada9eda8036b01fc6edc118a466e43cf04cbc8105a4a0306dc2a80b4091f

Observation 368fc9a0-32e7-48e7-81f4-69853d82ec85 · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Manning, Andrew Ng, and Christopher Potts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.664409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.664409Z digest=sha256:e3753286d9d11d81384958d5d056b4a965710debff1f106cab79e2563420d207

Observation 6b4450c4-38db-4f31-b009-ecea4e70e44d · outbound

This paper cites an unresolved cited work.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:32:04.720488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.675879Z digest=sha256:0506701e33f93a8f8fd39853a414dacd51a6c5d55987143294fa3df69e6e0f56

Observation 6e7cf706-c3cc-450b-9e5a-57ffaaadb444 · outbound

This paper cites an unresolved cited work.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:32:04.700027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.681799Z digest=sha256:9a138e5ad544c293ab4f61369e9956259a2cf3679698a821f8782e0f6d14ad2a

Observation a72df6b2-87ca-4b4b-a0f8-eb6ad8848330 · outbound

This paper cites Squad: 100,000+ questions for machine com- prehension of text.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Squad: 100,000+ questions for machine com- prehension of text

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.681840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.687831Z digest=sha256:93b016f3f274778c3184602558c6ecccce52b7823986d77ad681ddaa3048687a

Observation f607e9a1-5cd4-4e2f-84d4-3f398bb577c9 · outbound

This paper cites The pascal recognising textual entailment challenge.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts The pascal recognising textual entailment challenge

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:32:04.662745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T16:32:03.693730Z digest=sha256:053bf2a20909c9b27f0960ca7f0dc3869681e5030a645552e41a561bd8f44bce

Observation 8f91b630-358f-4061-8801-cb20e554f439 · outbound

This paper cites PaLM 2 Technical Report.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts PaLM 2 Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.698662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.698662Z digest=sha256:cdfdc248200b865b8dc49d5f6a3fd5a0fd5d02f687fa77d4a86c8fe4fbf81420

Observation 8a9d747b-9a42-4583-9d1b-02b07b4e6363 · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.704458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.704458Z digest=sha256:1960d9d9c2fd96c1e7c7a96cf9897afa6522c3d0bbc3932d69a401d26bf7289f

Observation fc30f7b7-3e5a-4e4a-80ed-637a98deb910 · outbound

This paper cites Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Language models are unsupervised multitask learners.OpenAI blog, 1(8):9, 2019

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.714918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.714918Z digest=sha256:90546be107025fdf0293c544c771004a9263bf92698c59bac6bd26f38d3c6e01

Observation 440c65a9-df59-4eaa-99f9-ad67e76f2045 · outbound

This paper cites FNet: Mixing Tokens with Fourier Transforms.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts FNet: Mixing Tokens with Fourier Transforms

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.720600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.720600Z digest=sha256:8a934235f7603298ed1e0145c5956c9d5b5219e421cb8f250f7f6995a8023877

Observation 598c47b0-3365-4852-a613-acafa17a55f4 · outbound

This paper cites Densely Connected Convolutional Networks.

Dynamic Adaptive Shared Experts with Grouped Multi-Head Attention Mixture of Experts Densely Connected Convolutional Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-15T16:32:03.727373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:32:03.727373Z digest=sha256:45cbf49e388befc74eb41cb351da2d2a609b877961daf6952084073ae177bb9b

Pith citing papers

No inbound Pith citation observations are available.