Pith. sign in

Paper Citation Record · LEDGER

SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2312.07987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.07987 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T19:40:09.937100Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:27:46.660519Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e45bede9-6c9d-4f26-ad44-2b34d4e07e10 · inbound

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective cites this paper.

Sigmoid Self-Attention has Lower Sample Complexity than Softmax Self-Attention: A Mixture-of-Experts Perspective SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:40:09.937100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:40:09.937100Z digest=sha256:06f3d6b58359db0cc8b64aa5ec8678e6ec23b5a2cf934df3814acbc445f93fc8

Observation 836bbf3b-d613-4e08-a16c-1ca30db11d10 · inbound

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection cites this paper.

Routing Mamba: Scaling State Space Models with Mixture-of-Experts Projection SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:43.909348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:43.909348Z digest=sha256:b92289f5fc5f20b23ae87fb7dc6fbeeaf4a9a8f7882fa20890c21cab7cbf2651

Observation f69c465c-4bd8-45d0-a382-94a3327a9334 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:21:28.252038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:19:12.696712Z digest=sha256:34f6eef48a58afb206c6adc7d92ca38bd101c95d7a5ee58e327d697b33ff8da3

Observation 82b6ef2c-655b-4686-ac88-54e268505ad6 · inbound

Sparse Layers are Critical to Scaling Looped Language Models cites this paper.

Sparse Layers are Critical to Scaling Looped Language Models SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:35:29.123697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:26:37.559459Z digest=sha256:4bb4a6237e8d6ea1f4ab7186f3a1147c48495d31911e4f803b12c94501f1df22

Observation 93ac1217-3f18-4a5a-8422-1ebd631c2aa1 · inbound

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation cites this paper.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.677614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4c533863e8755bcde9a897c74dd0bcc2d0722283817c1b96d100c22668ad8dd6