Pith. sign in

Paper Citation Record · LEDGER

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

As of 23 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2412.05540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05540 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:41:21.642571Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:43:25.901211Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:51:38.997059Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc710e34-59b4-4380-ab50-e3dab07a61f4 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.490482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.490482Z digest=sha256:67289fca1809a90c34281f6df99a74b92b0c9b6e2467626909610d67c333acc6

Observation e4b5cca8-1401-4f7e-b55f-14dfc5164697 · outbound

This paper cites Zero-shot text-to-image generation,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Zero-shot text-to-image generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.264997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.497115Z digest=sha256:e380365b8d7076fc66724f958c8f3ba168d4b5d9bf4d84a63669fe0e1e361b37

Observation a07b0ed2-e6b8-4132-8281-57948c57aafd · outbound

This paper cites Mistral 7B.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Mistral 7B

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.502618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.502618Z digest=sha256:3b0574c2b7fa4196e0c03f3adee78cc58f144a46a67f1ce10aa9baf6c03c5d6a

Observation 01231065-4c37-4995-a696-df65663d7198 · outbound

This paper cites Unified scaling laws for routed language models,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Unified scaling laws for routed language models,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.250930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.507634Z digest=sha256:4a1301124acd15236c0cc5ccbc8c0d30328b154dc4460763f6ac894bfe3044d0

Observation 630c04d3-6f97-4d36-8d79-ef099dfdd46c · outbound

This paper cites Scaling vision with sparse mixture of experts,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Scaling vision with sparse mixture of experts,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.519139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.519139Z digest=sha256:f78fbefc7b869fdabf93349cc3311ff007201bd98b4a6f8b8e77fa337d8a4e04

Observation b0469497-c796-44f9-9bbb-8f807b467062 · outbound

This paper cites M3vit: Mixture-of-experts vision transformer for efficient multi- task learning with model-accelerator co-design,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers M3vit: Mixture-of-experts vision transformer for efficient multi- task learning with model-accelerator co-design,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.226290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.526321Z digest=sha256:b4a9c87eca9e17a75645f57cd03674f67e479b1fa5c9995f79dfb093727a7feb

Observation 1dd75d63-800d-4c3b-a0e0-d562a164bcb6 · outbound

This paper cites Networks of spiking neurons: The third generation of neural network models,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Networks of spiking neurons: The third generation of neural network models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.207769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.534114Z digest=sha256:9bc98af609915d3615af99f9c7172c8349130c3f23db5fd4e1f41e380d1816ac

Observation d7eabd5c-0990-4a42-af51-2e010b8588c7 · outbound

This paper cites Towards spike-based machine intel- ligence with neuromorphic computing,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Towards spike-based machine intel- ligence with neuromorphic computing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.192790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.540861Z digest=sha256:9fb54dd5ef41eb7de26b8ca7de042205f43ff4d755773309ac39de75418bb3dc

Observation abc41cbd-dc55-4c6e-bd83-07b1e49cbee6 · outbound

This paper cites Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.179121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.546070Z digest=sha256:697cb9ab61c31550d5041de77d420c284bfbc8acfdd93180aaf643a7c17d8516

Observation bcc1ecce-cd5d-4e50-aed4-85d9e761055c · outbound

This paper cites Loihi: A neuromorphic manycore processor with on-chip learning,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Loihi: A neuromorphic manycore processor with on-chip learning,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.552958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.552958Z digest=sha256:a1163efb5eb47bb60210573904665d743a7edda45890f127267e50bf000c34f1

Observation 4f0bb142-1c41-4a8b-829e-7f5da7b0fb41 · outbound

This paper cites Spikformer: When spiking neural network meets transformer,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spikformer: When spiking neural network meets transformer,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.137091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.562208Z digest=sha256:6e176268c65059b730464006144d5f148fcfc5c78a1300ea175e4481e2d30470

Observation d3faa5c3-17da-4d16-a6eb-cf0ed2d29e29 · outbound

This paper cites Spiking transformers for event-based single object tracking,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spiking transformers for event-based single object tracking,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.113630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.566987Z digest=sha256:d3ea0c5e6b0680f13c7d0167f74ce643d290c9e2dad87c2abb9f9c2f7a3c9eb3

Observation c2a553b6-e50d-42ab-a6b9-3769f2de6c90 · outbound

This paper cites Spikegpt: Generative pre-trained language model with spiking neural networks,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spikegpt: Generative pre-trained language model with spiking neural networks,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.093632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.571540Z digest=sha256:140e46360cd3cd8ff6e2450dff956a0ef60489e635f0001f0529ab2cb9754192

Observation 77e862d1-3657-45d9-b92b-356ddae6465d · outbound

This paper cites Spike-driven Transformer.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spike-driven Transformer

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:21.933526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.576255Z digest=sha256:c727bc7aac46e28d097679beef14990aef5696e0843643d380b1848d699f9285

Observation c7930d23-617e-46df-8b11-b783b97b05eb · outbound

This paper cites Reconfigurable dataflow optimization for spatiotem- poral spiking neural computation on systolic array accelerators,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Reconfigurable dataflow optimization for spatiotem- poral spiking neural computation on systolic array accelerators,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.073022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.582056Z digest=sha256:660fd23b8328acc0e59089e453744bc2c0736281838d6fc8a6bd9e1515078751

Observation 734d311b-43f5-47c4-97cb-0b1646618ea7 · outbound

This paper cites Parallel time batching: Systolic- array acceleration of sparse spiking neural computation,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Parallel time batching: Systolic- array acceleration of sparse spiking neural computation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.153553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.587299Z digest=sha256:d9a888e557b1505f74dc85b17d2ba25b21fc34087b5b5afbe7072c5e5f8f36aa

Observation f117f8ed-ade5-48d3-86c0-bc289fc712bc · outbound

This paper cites LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers LoAS: Fully Temporal-Parallel Dataflow for Dual-Sparse Spiking Neural Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:41:21.910319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.592496Z digest=sha256:dba1b8df49f93fcb12d03517a7cc55553a12ed707bf033878cb6e36c918ba461

Observation 0496b4a5-050a-4289-a350-2dc3a3a3c674 · outbound

This paper cites Spinalflow: An architecture and dataflow tailored for spiking neural networks,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Spinalflow: An architecture and dataflow tailored for spiking neural networks,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.598783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.598783Z digest=sha256:a9dd6b15ed822a5daf23386fa3351d7036f7e8b333ca246e5e1e559bde4fc004

Observation 35c969a9-2f04-4a0a-a1b3-a4f7b879adf0 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.604890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.604890Z digest=sha256:23c0ca7f296f1c442608e4df2dbf7ea18f7f8f6363aae02344c959ee62318b34

Observation df4b7f45-241f-45f0-bb8d-3950f1efd3ad · outbound

This paper cites 3d-carbon: An analytical carbon modeling tool for 3d and 2.5 d integrated circuits,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers 3d-carbon: An analytical carbon modeling tool for 3d and 2.5 d integrated circuits,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.034996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.611982Z digest=sha256:599b70d90739be076d5d94c684bd09268995715353a36b7a438dd4e143a0321d

Observation 7031a519-f29e-4358-94ba-36e100f94e80 · outbound

This paper cites Nebula: A neuromorphic spin-based ultra-low power architecture for snns and anns,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Nebula: A neuromorphic spin-based ultra-low power architecture for snns and anns,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:22.014632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.616423Z digest=sha256:74bb093604d05438abbf47bd184eb870eb2fc542902c4d0a67ab2ee500189ddb

Observation 13cef4d3-103b-4b89-a464-67624381c2e9 · outbound

This paper cites 30.2 a 22nm 0.26 nw/synapse spike-driven spiking neural network processing unit using time-step-first dataflow and sparsity-adaptive in-memory computing,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers 30.2 a 22nm 0.26 nw/synapse spike-driven spiking neural network processing unit using time-step-first dataflow and sparsity-adaptive in-memory computing,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.995628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.621426Z digest=sha256:498ab4a1a481c6e86cf37f9e7f4e2a1a2d1bfc4b8f32aba2368fca22279dc4bd

Observation 55057ff2-f72a-4ab2-99c3-d6b7bc894bb0 · outbound

This paper cites Design and architectural co-optimization of monolithic 3d liquid state machine-based neuromorphic processor,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Design and architectural co-optimization of monolithic 3d liquid state machine-based neuromorphic processor,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.626932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.626932Z digest=sha256:79c16961f836393189c76b135d4bfb67801e37e3f50ca4d0884c926074dab067

Observation 20363ebb-9580-46ff-af03-082a1fca549a · outbound

This paper cites Area-efficient and low-power face-to-face-bonded 3d liquid state machine design,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Area-efficient and low-power face-to-face-bonded 3d liquid state machine design,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.979842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.631366Z digest=sha256:7e64cb9cefe2525cf1dfb1d35b33081e7cadcc9e5998cbbee21f4ac32bf322b4

Observation e1c9ebdf-6c5c-46f9-a8b0-d40daf0a8dc5 · outbound

This paper cites The cifar-10 dataset,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers The cifar-10 dataset,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:41:21.965393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-11T20:41:21.636773Z digest=sha256:adbeff490ddc38e754b525d17d9d56090c2f2c067860b7f1e72c638e1ad26742

Observation c91688db-3921-4659-9f99-8ec3f1d6f87c · outbound

This paper cites Pin-3d: a physical synthesis and post-layout optimization flow for heterogeneous monolithic 3d ics,.

Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers Pin-3d: a physical synthesis and post-layout optimization flow for heterogeneous monolithic 3d ics,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:21.642571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:41:21.642571Z digest=sha256:58a5b01bd223a5b31b5d3af8a80cfc64ba599344d65e16af3d1ad0c8a9d3d63e

Pith citing papers

Observation 8aa85490-e6c9-4ede-adb9-6bb8a5c07066 · inbound

SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks cites this paper.

SpikeX: Exploring Accelerator Architecture and Network-Hardware Co-Optimization for Sparse Spiking Neural Networks Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:43:25.901211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:43:25.901211Z digest=sha256:c2e9db9b606be0d3f88fc4bc02a659d38f501985c354cab6f762a4c2095ffd75

Observation 1a9c8c84-c348-4d5d-8571-33d990d5e2c5 · inbound

Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts cites this paper.

Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts Towards 3D Acceleration for low-power Mixture-of-Experts and Multi-Head Attention Spiking Transformers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:51:39.113913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T19:10:18.888463Z digest=sha256:013d0d7b7c451909b92e066b884f5240380bc6a31103890561f40c8d3fd3a4d8