Pith. sign in

Paper Citation Record · LEDGER

Mixture of Hidden-Dimensions Transformer

As of 22 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.05644.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05644 v3

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:38:23.762580Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:55:56.498566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T05:55:57.211220Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83395ac3-5424-4702-b68b-4f7a130ce83a · outbound

This paper cites write newline.

Mixture of Hidden-Dimensions Transformer write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.408668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.408668Z digest=sha256:63dff137913a5fe443d5a9c681557c87f994ec1a8cf356d92fceebccfb10266c

Observation 0f94a630-6bbb-421f-aaf1-0c5ef7562a01 · outbound

This paper cites LLM in a flash: Efficient Large Language Model Inference with Limited Memory.

Mixture of Hidden-Dimensions Transformer LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.416957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.416957Z digest=sha256:8f5e2c40ff36a5b74b51ae85fd3e9ba345eb9e5704c6717534952388c61be12a

Observation 8f8b5925-9250-40e9-a66e-2a7ff1573aa9 · outbound

This paper cites Anthropic: Introducing claude 2.1, 2023.

Mixture of Hidden-Dimensions Transformer Anthropic: Introducing claude 2.1, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.689837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.424840Z digest=sha256:745732eb8933021d18b566d5daa813bb1096a4b34a9a13b692b288c23cc3d624

Observation e8a30ba0-0fba-41e0-8832-aa881778c5e2 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Mixture of Hidden-Dimensions Transformer Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.430861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.430861Z digest=sha256:aecc6bf2419cf069bc1201ef6ac20e98f99eef15dc5b2f4e05cd97f14ff54220

Observation a8e3bab2-063e-47d6-8fed-297395bbb84e · outbound

This paper cites Flextron: Many-in-One Flexible Large Language Model.

Mixture of Hidden-Dimensions Transformer Flextron: Many-in-One Flexible Large Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.439913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.439913Z digest=sha256:3796e65b2e4a3f73f83a781f593361a68589c5fef3d874d16d476b734bb7ed9d

Observation f15ff33b-c536-4ea6-82f9-8afc2ae39bc5 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Mixture of Hidden-Dimensions Transformer A Survey on Mixture of Experts in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.448370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.448370Z digest=sha256:f30122e8c33312ac18f7ecfbc3477ba4d66538c6c2c2a1c04ab7204228c189fc

Observation 37d53d1e-ed64-402b-b827-b720d850a98b · outbound

This paper cites LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery.

Mixture of Hidden-Dimensions Transformer LoRAShear: Efficient Large Language Model Structured Pruning and Knowledge Recovery

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.456219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.456219Z digest=sha256:ad88df27718bec0199952937a70ea7b197b7fda7a7a86f24174f3f331240f7a5

Observation ba982374-9a63-42da-be62-79734065d780 · outbound

This paper cites LEMON : Reviving stronger and smaller LM s from larger LM s with linear parameter fusion.

Mixture of Hidden-Dimensions Transformer LEMON : Reviving stronger and smaller LM s from larger LM s with linear parameter fusion

Reference 8

Resolution
verified exact
doi, observed 2026-08-11T20:38:23.869746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.463461Z digest=sha256:e8130f48db2dfbfde8db1e1b1c9824f8e43f7402e1966f41192fcfa7a5c25933

Observation 3db74036-d71d-4dd3-a25d-58f511bfc2b0 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Mixture of Hidden-Dimensions Transformer BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.469365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.469365Z digest=sha256:13d24852bda57924f274840222c63a57f4532ea0cd056055d880ff002c778842

Observation 0bbd5d0d-dc7e-4f4d-be3c-b2064fa9ea15 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Mixture of Hidden-Dimensions Transformer Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.483149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.483149Z digest=sha256:39003cdb8e15c228feb54b0be9a9d0e204531d532cc6e56f904e9a99a4365714

Observation 964503fb-5db9-4afa-af69-5629f854c27f · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

Mixture of Hidden-Dimensions Transformer DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.492942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.492942Z digest=sha256:60c4cc2908b9a10265a6c6a39bb0e5f498f2bf39b80ca321dc96e085fb459629

Observation 36cc006f-df35-424a-bbe6-073d28722ee4 · outbound

This paper cites Monarch: Expressive Structured Matrices for Efficient and Accurate Training.

Mixture of Hidden-Dimensions Transformer Monarch: Expressive Structured Matrices for Efficient and Accurate Training

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.499076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.499076Z digest=sha256:884d01a452e23338a6121d9f298d6bd14169c94f85c642ae1d1a0616c2369492

Observation 251f4a72-c186-4a83-a231-be02c949b287 · outbound

This paper cites Introducing pathways: A next-generation ai architecture.

Mixture of Hidden-Dimensions Transformer Introducing pathways: A next-generation ai architecture

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.657866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.507762Z digest=sha256:484a577f9c80d32cf8858f613da888c654db893607c0b3123d8b5659a1885bd2

Observation 134d4614-fcdb-4d9e-8495-105c5b566cac · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Mixture of Hidden-Dimensions Transformer Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.514082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.514082Z digest=sha256:767c6b869680a47c975c989c38c6e3b1184f302c973fec8ca3e3fe31fc56a442

Observation a7c83a93-222a-4d18-9a22-83496f7fc5a8 · outbound

This paper cites A framework for few-shot language model evaluation.

Mixture of Hidden-Dimensions Transformer A framework for few-shot language model evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.521794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.521794Z digest=sha256:aea558875c4f92ecf07ec4deb82bd67e8b6f12076f44b27e0be1871aa51bb030

Observation 2e15dfd2-0f4f-43be-9971-1778a7965906 · outbound

This paper cites Mixture of A Million Experts.

Mixture of Hidden-Dimensions Transformer Mixture of A Million Experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.528934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.528934Z digest=sha256:27eaf713cd78c2bdddd1bad1c29254fda2a5589719d657105cf72f0cfaff3016

Observation 8b7687f0-0192-48a2-a4b4-0ab6c2483d55 · outbound

This paper cites Mixture of Nested Experts: Adaptive Processing of Visual Tokens.

Mixture of Hidden-Dimensions Transformer Mixture of Nested Experts: Adaptive Processing of Visual Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.534897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.534897Z digest=sha256:6153649a22227811ba0942d3be3f64fcf2cd76138930cd3e19dcb08ff7f6d2b0

Observation a4fd6eab-201d-4aed-9c98-c7c4935e0d0d · outbound

This paper cites Mixtral of Experts.

Mixture of Hidden-Dimensions Transformer Mixtral of Experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.547293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.547293Z digest=sha256:4d5cf48d81f1413f99e047e0ccc49792b7be42a8f7427c1f469b5c6ade721721

Observation fa8d3155-7b72-47c2-bec2-d53ff50e70a0 · outbound

This paper cites Scaling Laws for Neural Language Models.

Mixture of Hidden-Dimensions Transformer Scaling Laws for Neural Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.552691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.552691Z digest=sha256:c95222552d171cc9f19e745111b874791237027f0af59ee395811fdd481012f8

Observation 813f5a5b-7772-49da-be1f-3f3e0697d419 · outbound

This paper cites Distributionally Robust Receive Combining.

Mixture of Hidden-Dimensions Transformer Distributionally Robust Receive Combining

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.558130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.558130Z digest=sha256:0dcebc5d967df1f08c74257ead896d4f3f1e26d7760c5b54355c863971aaa044

Observation d6160df5-f70d-4b0a-ae11-44e9c8b9f3db · outbound

This paper cites Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models.

Mixture of Hidden-Dimensions Transformer Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.565308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.565308Z digest=sha256:ed817eb24b89d7d039334f12b14fed3f349f3b8d73e3e79ed28e47e695cfdc4a

Observation 0026c00e-54a2-4099-baa1-eeb290ce74ce · outbound

This paper cites LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning.

Mixture of Hidden-Dimensions Transformer LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.581695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.581695Z digest=sha256:e94ad4efaa51c8de2f51c6cffa9f178ec0cc3bd69abfa021ae72e128806e8f24

Observation c9d39f9a-2862-4fbb-bb56-cb5a9878fac3 · outbound

This paper cites Training-Free Activation Sparsity in Large Language Models.

Mixture of Hidden-Dimensions Transformer Training-Free Activation Sparsity in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.587766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.587766Z digest=sha256:2960f4582a2ff774d003de12c011d84a8e8c05ed96a07343585b69090c47daed

Observation acd7c951-1ff5-4c8d-bab9-564b570a548b · outbound

This paper cites Deja vu: Contextual sparsity for efficient LLM s at inference time.

Mixture of Hidden-Dimensions Transformer Deja vu: Contextual sparsity for efficient LLM s at inference time

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.628717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.595139Z digest=sha256:c306a109f2d65803a3b49af469103d3703dd2a80dc1780297380f28296b89fcf

Observation deb19e93-f95b-456c-ad2a-2e0a1970e471 · outbound

This paper cites Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time.

Mixture of Hidden-Dimensions Transformer Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.600708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.600708Z digest=sha256:fab0acdaa71b581f57b1304cf6c68020c846cae89c598106241ccace819eb248

Observation fab095ce-f06a-476b-bf6d-35a2bc12ecd0 · outbound

This paper cites LLM-Pruner: On the Structural Pruning of Large Language Models.

Mixture of Hidden-Dimensions Transformer LLM-Pruner: On the Structural Pruning of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.608435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.608435Z digest=sha256:90f3024fd06110543e1d8df3da3d4d4fdd98c98f38a7e5e38c9422d884eaa800

Observation efe7dd8d-2724-4841-bcb7-89752714703c · outbound

This paper cites ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models.

Mixture of Hidden-Dimensions Transformer ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.615270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.615270Z digest=sha256:1f2fe9d42bb21a072db42b20451e28ea0bbfd7ad70d9ab93d1456889837f3544

Observation 63bedd7a-4c85-4f41-b5b4-fb48275abb05 · outbound

This paper cites Openai: Gpt-4, 2023.

Mixture of Hidden-Dimensions Transformer Openai: Gpt-4, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.612221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.621670Z digest=sha256:557133971f39fef33dc6c14e2f2d82a0f13c649226f53e3ba65109a45976cebd

Observation 777f653c-fadd-41ed-98dc-0ea87a951471 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Mixture of Hidden-Dimensions Transformer The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.627060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.627060Z digest=sha256:85f05943de761cb57784087ef6333d1773e519b872cc3af4da7c492891959306

Observation 900cd83a-d1e2-4b7c-a866-5afdd78232d1 · outbound

This paper cites Unlocking emergent modularity in large language models.

Mixture of Hidden-Dimensions Transformer Unlocking emergent modularity in large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.633754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.633754Z digest=sha256:2584a31525729c127b624575f759678946189b5a772b4531d0c13fbc36b13fc0

Observation 7e99c9f0-680f-4d21-9f80-eb79e66ec93d · outbound

This paper cites Scaling vision with sparse mixture of experts.

Mixture of Hidden-Dimensions Transformer Scaling vision with sparse mixture of experts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.595141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.640743Z digest=sha256:cae7436b64a652b771d79b1fa930c1b90c86990cafe0b45ea0e35b180f916774

Observation 79213a82-fe64-45f9-b6ef-86747e80cee7 · outbound

This paper cites L., Bhagavatula, C., and Choi, Y.

Mixture of Hidden-Dimensions Transformer L., Bhagavatula, C., and Choi, Y

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.646709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.646709Z digest=sha256:059e46736f6ace5c448a974b96b813cffd7eea8b19c8df259837ca6104e9eaaa

Observation 4c9521c7-44c7-4ce4-9a29-0817068b233b · outbound

This paper cites an unresolved cited work.

Mixture of Hidden-Dimensions Transformer Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:38:24.572914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.651912Z digest=sha256:efdebf86ed7e3dcc90b33c51e53e325176da3b400db4aafcf7190d7b7ceb0759

Observation e18e27b0-5b0e-4473-8279-26e4045212f3 · outbound

This paper cites ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models.

Mixture of Hidden-Dimensions Transformer ProSparse: Introducing and Enhancing Intrinsic Activation Sparsity within Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.656929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.656929Z digest=sha256:828611a3dfb21c6c9e6b4c7afa3c78464305cc14b993f28210d91999cfb63853

Observation 5e44806e-2abe-484e-a7ca-7cfb77db5dcc · outbound

This paper cites PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU.

Mixture of Hidden-Dimensions Transformer PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.668636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.668636Z digest=sha256:c10762e182efc09d35f3eb41dccc8a874f74c213622c131dc2de38b49277f77a

Observation 6e121a36-00bb-4b59-b0c4-2d3a319431c3 · outbound

This paper cites an unresolved cited work.

Mixture of Hidden-Dimensions Transformer Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T20:38:24.550962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.673500Z digest=sha256:a3fbd4ff2623f8367d520db155af1532836324f887ed87541d3b0c67c4d1b307

Observation 43d8b85a-5d62-4c95-8e16-da0010f231b7 · outbound

This paper cites Redpajama: An open source recipe to reproduce llama training dataset, 2023.

Mixture of Hidden-Dimensions Transformer Redpajama: An open source recipe to reproduce llama training dataset, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.530943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.681227Z digest=sha256:18d92463337623bd894541f6f8862bdd846549d9c4c72280c4a88e49d2e18c49

Observation a30b4891-0793-42ce-98c7-0280b198a547 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Mixture of Hidden-Dimensions Transformer Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.694311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.694311Z digest=sha256:e3fb8bafcdc49551b16a779aff0864fe2f6f73a2d75afe76893590f864d4dcf1

Observation 2ba27e98-9d08-4492-8f82-c7a654853007 · outbound

This paper cites Attention Is All You Need.

Mixture of Hidden-Dimensions Transformer Attention Is All You Need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.700157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.700157Z digest=sha256:fcb5f6ffe858aea304e1f31117e92f4269c0a3018203d26ccdbb60ebbef733ea

Observation 67aaca1c-7019-4060-9e12-9535dd2ec51d · outbound

This paper cites Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation.

Mixture of Hidden-Dimensions Transformer Ladder: Enabling efficient low-precision deep learning computing through hardware-aware tensor transformation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.500933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.706684Z digest=sha256:0aa55aa960b565665c2c721c8668449f8c9680632735a2541426e6ada365b3e5

Observation f408b0e1-b86c-4963-ab88-ffa3c3b63ba3 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

Mixture of Hidden-Dimensions Transformer Crowdsourcing Multiple Choice Science Questions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.712606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.712606Z digest=sha256:12ccda2b24cf3954054358496545dd5e22e7709e7d6ca51c3b88ad5a0884e45f

Observation 75637d93-5dd6-4d99-9a52-a2304ee1efc5 · outbound

This paper cites Multi-Head Mixture-of-Experts.

Mixture of Hidden-Dimensions Transformer Multi-Head Mixture-of-Experts

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.718220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.718220Z digest=sha256:c495589469128bfe4a798dc01ebf59d3e299f977f45306d3c58644729d695a10

Observation 7e6b0d3f-bbd4-44ab-9983-9641874246e1 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Mixture of Hidden-Dimensions Transformer Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.724341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.724341Z digest=sha256:69a6782c9719bb216119bbc3927fa82fdeb08c545278fd00b7e29c76630fee9c

Observation 3c8572ba-76b0-4d17-8bee-856fca40d0e2 · outbound

This paper cites OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models.

Mixture of Hidden-Dimensions Transformer OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.738488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.738488Z digest=sha256:13f14ee994d7c49d82679456fd0bab09cb9243dc69624933186d56142adc5377

Observation 1f73b69a-b2e1-49d4-bc5a-c7312741f42e · outbound

This paper cites The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers.

Mixture of Hidden-Dimensions Transformer The Lazy Neuron Phenomenon: On Emergence of Activation Sparsity in Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.744646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.744646Z digest=sha256:5fb733f49a79769de05fbc33811431389ee976c49b20410d3d6200d7cbf277b4

Observation a3372cce-7bf2-47dd-bf62-0e08fccbba74 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D.

Mixture of Hidden-Dimensions Transformer Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.751649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.751649Z digest=sha256:e57c6e426472b8dbbbd095646dd9603ba38d244c335cd39c9b67136f2320edb9

Observation e4b8fc5f-2351-4276-a861-d726c0ea1ede · outbound

This paper cites ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs.

Mixture of Hidden-Dimensions Transformer ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:38:23.757009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:38:23.757009Z digest=sha256:ccdd1b93e656e4eac25be52867fbcbef282981d9a2ca53f3b3e2c7877ee2fdca

Observation 900fa513-18e7-4f17-9125-4cbef946aa94 · outbound

This paper cites M., Le, Q.

Mixture of Hidden-Dimensions Transformer M., Le, Q

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:38:24.480949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-11T20:38:23.762580Z digest=sha256:c5dd9d2fb38adc7e0fffc5459b64e964d5f3c5e651459a872f8b12519d29db09

Pith citing papers

Observation b8508283-0f69-4d8c-bc68-e18a0f13dc5f · inbound

Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models cites this paper.

Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models Mixture of Hidden-Dimensions Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:55:57.215318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:55:56.498566Z digest=sha256:0aa18f72583ebe01b9da83c9b6531c85471cfa7e5731d52640768a746714e495