Pith. sign in

Paper Citation Record · LEDGER

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.15414.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15414 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:24:00.873308Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T01:39:29.681621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T03:24:12.966811Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved11
  • parse uncertain1
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13d87c56-375b-46a8-8a0d-346442c8da75 · outbound

This paper cites Cnn mixture- of-depths.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Cnn mixture- of-depths

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:06.069066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.027851Z digest=sha256:c744e400eaa17f6a16bd8f653bd4cf0c448cd0bcd22734cf8dbdabaa31ae017f

Observation 5955cf6d-ed92-42c0-96af-05ca9fa94802 · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:24:05.984198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.062940Z digest=sha256:e85e858dfe505c77a1d40e66352b5b964f0f161a769d94c4531e0fc1c145a246

Observation 96f4ab52-d866-4f7b-aac4-fa6b936dd4da · outbound

This paper cites Adamv-moe: Adaptive multi-task vision mixture-of- experts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Adamv-moe: Adaptive multi-task vision mixture-of- experts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:05.917478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.096642Z digest=sha256:893783fef0d92d98601e02d5963039bf5ab6a8c103af03b09be5465e44fd6f18

Observation dc4002fa-a2bd-40fc-8f01-c8a8488f97d7 · outbound

This paper cites Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Mobile V-MoEs: Scaling Down Vision Transformers via Sparse Mixture-of-Experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:58.172551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:58.172551Z digest=sha256:c7c5c1706dd6ae66ace711fb055517c6458b5b380380372a32fa3651a779d02b

Observation 6e5d2e1e-385e-4e1e-a60a-c4da269eb563 · outbound

This paper cites Editing factual knowledge in language models.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Editing factual knowledge in language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:05.709176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.271023Z digest=sha256:7933d29fedf05d4d335256d6971898e5b1d7135236733c85f62c4495257781bb

Observation ccb394a9-2ec0-444e-9ff6-25bed08d6f25 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Imagenet: A large-scale hierarchical image database

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:58.313312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:58.313312Z digest=sha256:70a96eef408e9191a86e3c3bfd21247c860f8df7c526b001ce24ba1878617e79

Observation 66ace745-8c51-4312-8e1f-c16018d37124 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks An image is worth 16x16 words: Transformers for image recognition at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:58.365441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:58.365441Z digest=sha256:28b811201e43d85aafe3ed657bef65d03b3b9ea0b9c72ab88f7a43df0f14fd9e

Observation ce600041-3fcf-4e20-ae9b-c6a3bd9c6c31 · outbound

This paper cites Learning factored representations in a deep mixture of ex- perts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Learning factored representations in a deep mixture of ex- perts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:05.255190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.514743Z digest=sha256:0cdbeddb2651a7bfa5d668d153bd37f9c66f25f6f4d6c2a3ce8364f45a5476ef

Observation 631009a8-1678-45f5-8cd9-fc1e32b6ee23 · outbound

This paper cites Switch transformers: scaling to trillion parameter models with sim- ple and efficient sparsity.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Switch transformers: scaling to trillion parameter models with sim- ple and efficient sparsity

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:05.017008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.570505Z digest=sha256:6ada1e911cbaf3b9377d6bc10b23c816c5134c984f5b8b5cdc5356b81ce7d28c

Observation 904a807d-9fad-443c-b007-efd9c1b9af11 · outbound

This paper cites Transformer feed-forward layers are key-value memories.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Transformer feed-forward layers are key-value memories

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.772928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.615645Z digest=sha256:2a52d0b4b66f405d982d21fb3b2c314d0a88af89a61b8a4bc6383629b339ea09

Observation 8917cfb3-4389-420f-8abf-d3aeb1e354d5 · outbound

This paper cites Paca-vit: Learning patch-to- cluster attention in vision transformers.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Paca-vit: Learning patch-to- cluster attention in vision transformers

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.669288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.662240Z digest=sha256:ede3cfec61768a9f0d26f17a3dee09c5c74ea231e8ac28d5c95fc97299e5bbfd

Observation 8964dc15-d008-444b-895d-4638759fe8cd · outbound

This paper cites Jacobs, Michael I.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Jacobs, Michael I

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.563268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.733747Z digest=sha256:6fc25e397d7bd1b513d7358f8b245b9e8fb22bd2d045584463abce8ce7305190

Observation c0cca5f0-338b-4d66-b4ab-4b96b2d64bf8 · outbound

This paper cites Learning multiple layers of features from tiny images.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Learning multiple layers of features from tiny images

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.461482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.835071Z digest=sha256:60d414a0e011e8d9e06160d1feff525bbb4f39066cdc184327635ddc4b10f988

Observation fdbd8935-4465-4e1a-b3e0-07f5e5f785d2 · outbound

This paper cites Gshard: Scaling giant models with conditional computation and automatic sharding.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Gshard: Scaling giant models with conditional computation and automatic sharding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.355795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.909804Z digest=sha256:6b96383da341756954a3989f3211435acd0b8bb9e7d1e3bfa38349ba721ae08d

Observation 8d1fdac2-013a-4550-a0f4-5c03088be907 · outbound

This paper cites M$^3$ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks M$^3$ViT: Mixture-of-Experts Vision Transformer for Efficient Multi-task Learning with Model-Accelerator Co-design

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:24:01.043720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.941561Z digest=sha256:da2ac23e09147c83424107efebf44ae8dd418c1705719859298eba3e4eabeb39

Observation 1257fd36-0f29-431a-919b-32b6b706072a · outbound

This paper cites Liang, Yiming Cui, Qifan Wang, Tong Geng, Wen- guan Wang, and Dongfang Liu.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Liang, Yiming Cui, Qifan Wang, Tong Geng, Wen- guan Wang, and Dongfang Liu

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.249861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.990130Z digest=sha256:84b8602f748a4083a6ef455034f87ef3b32aef6b3b09e51efb17f3cbd3020281

Observation 12356daa-86e5-4abf-85e6-88b94ce3a4e5 · outbound

This paper cites Expediting large-scale vision transformer for dense predic- tion without fine-tuning.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Expediting large-scale vision transformer for dense predic- tion without fine-tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.163106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.088942Z digest=sha256:b9c4e7cc658ff56e6c534fbd077152484d8f96ffd604d8a6d4e92d32d7176004

Observation 378c6db7-499f-468c-8b75-a6a4797b9128 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Swin transformer: Hierarchical vision transformer using shifted windows

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:04.060206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.169834Z digest=sha256:4a33eb1d9cfd5897fe375656a044b7003c00ec5a832c085530639d3b76c5cb1d

Observation cfeab651-0f64-484d-bf52-138c75035197 · outbound

This paper cites A ConvNet for the 2020s.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks A ConvNet for the 2020s

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.916903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.216226Z digest=sha256:9471e59e5877ee86551f322bd3589add017e288af5a54aecd1faa7f20c2765ab

Observation 14c2299d-7fc3-4264-83b6-6e543c39779d · outbound

This paper cites SGDR: stochastic gradient descent with warm restarts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks SGDR: stochastic gradient descent with warm restarts

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.672553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.277499Z digest=sha256:04f4fcc0922533de82adbcd7e9781cade4301b6e0e74db2b77f1ed3f51defa2f

Observation 0b6889e2-02a3-4894-84e1-f53876c982d0 · outbound

This paper cites Decoupled weight decay regularization.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Decoupled weight decay regularization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:59.331213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:59.331213Z digest=sha256:d8dea20535de0360d958266c5308bcd23d824d9a93fafccdbf835936847d584f

Observation 6c3a8ffd-db7a-465b-8cb4-05de588b739d · outbound

This paper cites Data-free dynamic compression of cnns for tractable effi- ciency.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Data-free dynamic compression of cnns for tractable effi- ciency

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.542747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.401092Z digest=sha256:873c84364f9dfb6ff222f5e81748c1b627d6678eee22f19e8e59a7c1676a609e

Observation 910002f6-248c-44df-840b-cc112a17533a · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:24:03.433425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.459275Z digest=sha256:e96ee14a3168d99c92689181cfb334c8ead298fefc616d3162e836b7bb1c4cd4

Observation 1924e943-e112-456a-9d42-9089d2f203b2 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.318844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.525792Z digest=sha256:406f9c12ea0d38921a9c176e4107f0f7105ea0cc773b5a88d0a6b9e9ea77803d

Observation 93298046-c3cc-4af2-a7fa-5028a97c44bf · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:59.618392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:59.618392Z digest=sha256:c4a78744280d18f653f93c0974a79fcf9298842416ea549f25f9e220d09f14fa

Observation 8d4740d5-ad37-43e8-81a3-099d565a82e3 · outbound

This paper cites Scaling vision with sparse mix- ture of experts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Scaling vision with sparse mix- ture of experts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.233111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.675519Z digest=sha256:6f10ac75b3158b12ae9dab377765a61c89a34844e73bd316d2b20b827561e4a4

Observation 29066870-2cf7-45e2-9ee2-4dfeaee96589 · outbound

This paper cites Le, Geoffrey E.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Le, Geoffrey E

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:03.091241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.734470Z digest=sha256:9ee9581e3320beff915bbb1106e38b718dc4590c47e22a07a739eb767e79203c

Observation 93ca7a46-5006-45ea-b77f-850162a1cd72 · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Training data-efficient image transformers & distillation through at- tention

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.997813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.795232Z digest=sha256:70d863bb3547bbfb458fcf14087e271fb74fccbbba04ce28216d48f6317d248d

Observation 4a86029b-a046-41b3-a71a-2c06788b8ef4 · outbound

This paper cites Gomez, Łukasz Kaiser, and Il- lia Polosukhin.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Gomez, Łukasz Kaiser, and Il- lia Polosukhin

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:59.844233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:59.844233Z digest=sha256:8a172ed6a39268b3516d3e322a9eb53d004ec1a63ccf3c82f7f4b0220bf4f0fe

Observation a4a5c7d5-5681-4b63-93e7-4cefdb72b6e5 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Evo-vit: Slow-fast token evolution for dynamic vision transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.847038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:59.974509Z digest=sha256:93d833577b453467cbea8d4807b3e8b4ec53491cf278ac02476cf3d9ce61531a

Observation 464e8b07-e0fd-432c-8be5-2bf496bf8ce9 · outbound

This paper cites Width & depth pruning for vision transform- ers.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Width & depth pruning for vision transform- ers

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.716047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.032714Z digest=sha256:ce7fd9bfb265567a0868bc6b639ce183a71c7056051031323450cbe08797efd3

Observation 46ccf3f7-4c0f-40f7-a545-e83710bc045c · outbound

This paper cites MoEfication: Transformer feed-forward layers are mixtures of experts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks MoEfication: Transformer feed-forward layers are mixtures of experts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.606178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.099921Z digest=sha256:8ee5d5417b53af7382eb578de179335dab581d1598a9413e04c9b9ceb1e9118e

Observation 3f4d5c2c-c54e-47ca-b10d-a5de3b9728c6 · outbound

This paper cites Emergent modularity in pre- trained transformers.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Emergent modularity in pre- trained transformers

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.476896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.175715Z digest=sha256:a27327ad774e1527d84d5bbecb4b849c48fee37d87e0dacaebaef4bdc9b70d49

Observation 4ecc2086-d3ea-4b26-92e3-d796a7e1ac49 · outbound

This paper cites Vision transformer pruning, 2021.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Vision transformer pruning, 2021

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:02.348371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.259021Z digest=sha256:977b821aac4ac1d1deb9aa243de9ab39ae6f0e51d35a025018ae9ce722f23119

Observation 3e76fd50-b792-4bb0-8f2a-d146be86147a · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:24:02.223252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.337829Z digest=sha256:25c8f2ef962285bd146acad9367dc600b44786a152067dc09ada846c7d318c75

Observation 4d79b088-41e4-4879-86f9-9e8cf99b9a26 · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:24:02.107505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.400571Z digest=sha256:a16466ddcffbfa31a016eac513de14d8d89cf05a62687db4442a4cc96bd006f0

Observation a314ddf4-c944-4124-9f95-dd19b142c509 · outbound

This paper cites [32] rely on weight co-activation graphs and a manually set number of experts.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks [32] rely on weight co-activation graphs and a manually set number of experts

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:01.989832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.469962Z digest=sha256:84f2fa2eba5a4b68b0c4bf534ac64b65ea68945c1d6083314bbe27319db9a0d9

Observation 522bc121-5af9-47d9-a820-fbe113f01fe4 · outbound

This paper cites Our results on ImageNet-1k validate this design, showing performance gains in vision tasks.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Our results on ImageNet-1k validate this design, showing performance gains in vision tasks

Reference 40

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:24:01.880385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.534943Z digest=sha256:a7cde00815360d096f2f57020db5d0207c06707e9265913d9b3e37a9c3b4e628

Observation 3b136eb2-b614-4b1f-ae4b-9a2d4acda1d6 · outbound

This paper cites Despite structural differences, all variants benefit from fewer MACs, reduced parameter counts and competitive final accuracies.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Despite structural differences, all variants benefit from fewer MACs, reduced parameter counts and competitive final accuracies

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:01.716240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.589673Z digest=sha256:f374b7a132b041d4dfa7ffdbe18eab1fa85fe701b8cda3addd48348adc0119a3

Observation 894ff46d-85d3-4a68-a614-7325debf04fa · outbound

This paper cites We compare our default method, HDBSCAN, against other density-based methods (DBSCAN, OPTICS) as well as partition-based alternatives (K-Means, BIRCH).

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks We compare our default method, HDBSCAN, against other density-based methods (DBSCAN, OPTICS) as well as partition-based alternatives (K-Means, BIRCH)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:01.595145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.681243Z digest=sha256:b70b46c60064c90d31c8526adb466fc8a377e8b40d10a00225f4fbbd5e955da8

Observation cb20b120-ee0d-45e7-810e-47765ebcac3e · outbound

This paper cites For the ex- traction strategy, we consider two methods for selecting the hidden neurons of each expert.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks For the ex- traction strategy, we consider two methods for selecting the hidden neurons of each expert

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:01.479245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.765634Z digest=sha256:1122be11626c6b0cc125d0c4b69811d88ad8068f08aa16d2a874fb801b41e6b9

Observation d9f42eb1-d16b-48a4-be7b-9e98b49b916f · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 44

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T15:24:01.344918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.830183Z digest=sha256:338c88543dd50fe20bdc043ec80b8fe58b6cac0a57b44df1bff75a8733da8f4a

Observation f8db9359-6281-4632-8542-2f31436fb03a · outbound

This paper cites Notably, the earlier lay- ers do not exhibit any formed clusters, reflecting the more general feature representations in shallower layers.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Notably, the earlier lay- ers do not exhibit any formed clusters, reflecting the more general feature representations in shallower layers

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:24:01.184009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:24:00.873308Z digest=sha256:bb14e075948d13eb2e7cfd07d544e9402d1bcb2bb958ccca300b1d956f7da045

Observation 28c7a30f-8255-4b0c-b108-e1bfc14aed30 · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 2017

Resolution
parse uncertain
no resolver link, observed 2026-08-07T15:23:59.912198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:59.912198Z digest=sha256:ffcba00fb50cc089e4160ac070f98e88c36ea12f57d10e89c3670fb16e996587

Observation 3e0dfc0a-34fe-44b9-a81d-761399d9620c · outbound

This paper cites an unresolved cited work.

Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks Unresolved cited work

Reference 2021

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:24:05.475124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T15:23:58.422257Z digest=sha256:dceb9faac5114eb3e8b05261644f103a22c74c44a0020da776eb8aa0a3b674b1

Pith citing papers

Observation d2073092-f490-4894-ba08-6d98b6206114 · inbound

CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer Selection cites this paper.

CLEAR-MoE: Shared-Basis Expert Extraction from Frozen Vision Transformers via Calibration-Driven Layer Selection Efficient Data Driven Mixture-of-Expert Extraction from Trained Networks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:24:12.968335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T01:39:29.681621Z digest=sha256:89f669fa4915cb0c9c66b17334a4fa227eb3c8657ba9b1995c1284c5a94be91d