Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:49:17.483517Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2412.04139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:49:17.483517Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:55:02.197434Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-09T12:08:23.021613Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 48ae1320-e6a8-437e-ac3a-4fed29d17f1b · outbound
Monet: Mixture of Monosemantic Experts for Transformers Nemotron-4 340B Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b05858b-9258-499a-bb95-2cdf3134d911 · outbound
Monet: Mixture of Monosemantic Experts for Transformers b", nn.initializers.zeros, b_shape) 11 12def __call__(self, x, g1, g2): 13x = nn.relu(self.u(x)) ** 2 14x = jnp.einsum(
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f52d3f37-a5a4-4a4a-976d-58e3e238cf44 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Haozhe Chen, Carl V ondrick, and Chengzhi Mao
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3bd332f-a45c-490c-81c8-a25fd29239de · outbound
Monet: Mixture of Monosemantic Experts for Transformers F**kyou!F**k (...)* (16.68%)(...)Snakesonamotherf*ckingplane
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cb4d296f-f8ed-4751-80d6-5f80c23a81c6 · outbound
Monet: Mixture of Monosemantic Experts for Transformers RealToxi- cityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 50b5e7a9-de52-495d-b6ee-ae6976f0760c · outbound
Monet: Mixture of Monosemantic Experts for Transformers Transformer Feed-Forward Layers Are Key-Value Memories
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 079ff9d5-2989-4396-b146-5c1b7bf9854e · outbound
Monet: Mixture of Monosemantic Experts for Transformers Mixture of A Million Experts
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7be31826-7cc4-468c-832e-0eaaf53d73fc · outbound
Monet: Mixture of Monosemantic Experts for Transformers An Overview of Catastrophic AI Risks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 282fbb42-89ca-4431-92a4-f850a48bdc8d · outbound
Monet: Mixture of Monosemantic Experts for Transformers AI Alignment: A Comprehensive Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8d396f3-f03a-4a46-8b79-f4fe4f2aae47 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Mixtral of Experts
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5837cd5c-81c5-4297-9768-8331a34d6c41 · outbound
Monet: Mixture of Monosemantic Experts for Transformers The Stack: 3 TB of permissively licensed source code
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c909ed0d-cae8-4be6-a2ea-195e35dfac81 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Theory on Mixture-of-Experts in Continual Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b0c3ed-ca01-4bd2-b729-05e8f79176b7 · outbound
Monet: Mixture of Monosemantic Experts for Transformers StarCoder: may the source be with you!
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f7b35d-b5d2-4d89-b75d-2b9c9eed1692 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Scaling Laws for Fine-Grained Mixture of Experts
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27bd9235-fb72-459e-ac69-e41d5c8db21b · outbound
Monet: Mixture of Monosemantic Experts for Transformers Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5e0d1f-8b6d-4135-a950-e7c1c5844321 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Jesse Mu and Jacob Andreas
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e650a4d8-dfe9-47c6-b120-9261da758d87 · outbound
Monet: Mixture of Monosemantic Experts for Transformers OLMoE: Open Mixture-of-Experts Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17cd2fc4-25f7-46dc-b552-0baf645a894d · outbound
Monet: Mixture of Monosemantic Experts for Transformers Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aefa3d9c-84a0-4219-83bf-2b3f5d689843 · outbound
Monet: Mixture of Monosemantic Experts for Transformers The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa7fe330-0eaf-4015-bae2-c6511e5a16a9 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Taking features out of superposition with sparse autoencoders
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8cb3b353-fb27-4997-a56f-98bdb79ee9c0 · outbound
Monet: Mixture of Monosemantic Experts for Transformers JetMoE: Reaching Llama2 Performance with 0.1M Dollars
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dc2d541-93b0-4641-9cde-bbe132a9a1b0 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Codebook Features: Sparse and Discrete Interpretability for Neural Networks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e7dff76-7ce8-4d30-b7f3-bfae6a2551ac · outbound
Monet: Mixture of Monosemantic Experts for Transformers Gemma 2: Improving Open Language Models at a Practical Size
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c0a478-58f3-4a12-b0dd-135ab96bcabe · outbound
Monet: Mixture of Monosemantic Experts for Transformers ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ddfaba-7cce-45f2-ab6d-bcace9afc643 · outbound
Monet: Mixture of Monosemantic Experts for Transformers CONTENTS A Method Descriptions 18 A.1 Expansion of Vertical Decomposition
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5787dd54-254f-43b2-bf12-6405cace9c38 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Moreover, the multi-head expert routing probabilities are consoli- dated into single routing coefficients PH h=1 ˆg1 hi and PH h=1 ˆg2 hj, reducing redundant aggregations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b521dc35-7029-4b51-819d-64fea07fada6 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 205d08a4-c02a-4e40-9b3f-dce2b2e89ee2 · outbound
Monet: Mixture of Monosemantic Experts for Transformers b1", nn.initializers.zeros, b_shape) 16self.b2 = self.param(
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2155c7b8-ccad-4b50-88cc-fd5497bd2339 · outbound
Monet: Mixture of Monosemantic Experts for Transformers To manage computational resources effectively, we adopt a group routing strategy wherein the routing probabilities are reused every 4 layers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b6d7aa3-ec52-4887-8481-d7afdbb26e34 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d3aad49d-4168-44ca-8d6c-53e51176b194 · outbound
Monet: Mixture of Monosemantic Experts for Transformers The ta- ble reports the number of experts assigned to each programming language across all routing groups
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0ad5e76-b44c-44e2-acc1-27ddb37e82c7 · outbound
Monet: Mixture of Monosemantic Experts for Transformers "" 12#!/usr/bin/env bash 13 14echo
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 248b5d53-cb14-42fb-a7cf-2921df6104fa · outbound
Monet: Mixture of Monosemantic Experts for Transformers Based on this, we masked experts associated with each language and re-evaluated the code generation benchmark to estimate the model’s capa- bility to unlearn programming languages
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ff23ca77-efa7-4ac1-896b-d436f02b129c · outbound
Monet: Mixture of Monosemantic Experts for Transformers JULYIV (...)rew (59.50%)(...)TheembroideryreadsinHebrew:
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5dfbb1be-ff90-4064-873d-92e3824017e8 · outbound
Monet: Mixture of Monosemantic Experts for Transformers BatchTopK: A Simple Improvement for TopK- SAEs.AI Alignment F orum,
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3be18078-dc49-4ee3-8d8c-5a5a606dd494 · outbound
Monet: Mixture of Monosemantic Experts for Transformers Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38440118-84ef-488e-9c34-51d3bc740853 · outbound
Monet: Mixture of Monosemantic Experts for Transformers A Review of Sparse Expert Models in Deep Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f391b43e-99f7-4df1-b0be-83ca33bcbf8c · outbound
Monet: Mixture of Monosemantic Experts for Transformers Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a2481438-38c2-41c0-b057-1993764de5aa · outbound
Monet: Mixture of Monosemantic Experts for Transformers Mem- ory Augmented Language Models through Mixture of Word Experts
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e3e18910-6f01-4337-ad66-3b0d06010f01 · inbound
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition Monet: Mixture of Monosemantic Experts for Transformers
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9802ffa-0cfa-4a99-b7e6-efa79ea705d1 · inbound
Studying Cross-cluster Modularity in Neural Networks Monet: Mixture of Monosemantic Experts for Transformers
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.