Pith. sign in

Paper Citation Record · LEDGER

ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2402.03804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.03804 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:37:39.507833Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:17:17.632979Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 287bf1e2-4352-430b-ac76-5ab02d64ee48 · inbound

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense cites this paper.

Activation Approximations Can Incur Safety Vulnerabilities Even in Aligned LLMs: Comprehensive Analysis and Defense ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:37:39.507833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:37:39.507833Z digest=sha256:9ba7a12e9c7ccf91a13e4e130a29ef74381fe6a3d3850524b5a957246cf50146

Observation 2f34fcf1-2e4b-4e23-b20d-f0b3a8f928a1 · inbound

Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity cites this paper.

Accelerating Linear Recurrent Neural Networks for the Edge with Unstructured Sparsity ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T15:48:57.426567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:48:57.426567Z digest=sha256:e6b54e75f21c955325eb5ca42e49287fb9e5f1becbd9bcb24aa0b4c626620151

Observation 40040ca2-e341-43ff-a3dc-123c505b6d54 · inbound

Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning cites this paper.

Why Neural Network Can Discover Symbolic Structures with Gradient-based Training: An Algebraic and Geometric Foundation for Neurosymbolic Reasoning ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.715111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.715111Z digest=sha256:9da2d07335831b4e05f166c5f67ad484dc97dd032b7896ae78ab50c48e3f4346

Observation 01cbe2a3-37ea-4fb4-8e35-e1dc8e9e3ef6 · inbound

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity cites this paper.

BlockFFN: Towards End-Side Acceleration-Friendly Mixture-of-Experts with Chunk-Level Activation Sparsity ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T18:20:00.484184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:20:00.484184Z digest=sha256:eb83fad188412d06e1a8523fb93ef50957c4eb1db5232eb3ce47696533107c84

Observation 85ee922e-f1ab-4c36-b5dd-a71cce40314d · inbound

SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment cites this paper.

SmallThinker: A Family of Efficient Large Language Models Natively Trained for Local Deployment ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:09:56.533540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:09:56.533540Z digest=sha256:c73c676283694f30b2d22b3b59d455ead0fdf77b1e904d2e8ebb7fed360b5a3f

Observation f7ba7e3f-7e6d-4472-ad5e-83c3cfdb7346 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:21:18.973013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:19:25.483640Z digest=sha256:a2fef5005c962fa60e67fea27a25c6f30974cb9af3040283dc83ecdb99cdc692

Observation ebd87e62-d818-459d-a6ee-8f1ba7a219b2 · inbound

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity cites this paper.

Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via Spontaneity ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:54:51.166503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T11:54:29.436149Z digest=sha256:5b9845e4c8f2f238064da36d4eb410113cbbd8dec1f47fd0e9862b6fb4613f5c

Observation bcd43197-16ec-429a-aaca-503f6635dea3 · inbound

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers cites this paper.

Physics of Language Models: Part 4.1, Architecture Design and the Magic of Canon Layers ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T15:22:56.382733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:22:56.382733Z digest=sha256:b9a48b622ae4ad7dd981b4afea7a5c958bc4c8d43486f7cbefd09c97ff1903aa

Observation 7077b6b3-3a94-4a55-89da-f01a581e644d · inbound

Parcae: Scaling Laws For Stable Looped Language Models cites this paper.

Parcae: Scaling Laws For Stable Looped Language Models ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:01.138815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:33:04.442462Z digest=sha256:612c22edf4ebe4c282666a3cbc4dcd106ce519f5d8116c2de268e4bb69fda3b2

Observation 3c8c1dfc-ee14-425f-a63b-91faf47b244f · inbound

PointTransformerX: Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms cites this paper.

PointTransformerX: Portable and Efficient 3D Point Cloud Processing without Sparse Algorithms ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:18.093343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:43:54.832085Z digest=sha256:d4401f01b2f8c2544b70b84bf369f96de1ac339988b7c6cf9cf9251cb30cc95b

Observation b9361ec4-799f-40f6-848d-6f9a5cacc4d2 · inbound

The Role of Symmetry in Optimizing Overparameterized Networks cites this paper.

The Role of Symmetry in Optimizing Overparameterized Networks ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.426128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-07T16:53:30.229571Z digest=sha256:132dd30a1d7c6aa2b8b236f541e1000ca1699b12319f3f20b98cfe15b8d4b3f8

Observation 2639d1f2-48c3-448b-a1e9-6b404ae26151 · inbound

The Role of Symmetry in Optimizing Overparameterized Networks cites this paper.

The Role of Symmetry in Optimizing Overparameterized Networks ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:58.620267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T00:59:12.261460Z digest=sha256:2102f7e9a078a67cecd6c93db5147d532e68deb5336c896b8ba9284b8ee65da9

Observation 30a0bd63-8769-4429-bb5c-4526375eb570 · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T13:48:19.644508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T13:46:32.405079Z digest=sha256:26691532b14280dbcea9559ff622a987536c8e940da3a40403375a652a4f8169

Observation b4ebf4f8-f677-4de6-92e1-4ceec936a723 · inbound

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes cites this paper.

Bug or Feature$^2$: Weight Drift, Activation Sparsity and Spikes ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T09:16:19.699136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T09:15:32.395442Z digest=sha256:e501ff9e871c2a3eb94da4e535c7c6be23d251700f53ecee5195f9058e8ea0ec

Observation b92b3a43-0820-419f-af66-e9f00abc0dc5 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.331701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:022d478b164c56609a3f978a2bfe28d95a693d8855ddc72c237f089927a35270

Observation 50288ef7-acbf-4f4e-a6f6-726120bd4fb1 · inbound

RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models cites this paper.

RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Models ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.728381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T19:40:42.033793Z digest=sha256:741852678d644c53603a5c7baa4b38fe3267679158dbe6ce4174e56bfc9cf3f7

Observation f7dfc51b-69b5-4ae7-82f6-142080e4be9a · inbound

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations cites this paper.

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:54.835580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T19:37:51.563121Z digest=sha256:f860b62b858adca1a02ad1d074946b4f90e9e15a6c789fcce49b3f1e5f617464

Observation 651bf497-87b6-4403-9d40-04fe69698b12 · inbound

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination cites this paper.

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:17:17.635430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T19:16:31.055024Z digest=sha256:b96c2f4aeacf27bbb1526df53b3a76ed65642eec25a32c910217aeddabf2b0b3

Observation 7879e60c-da60-439b-ad24-41cebf40c861 · inbound

SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels cites this paper.

SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels ReLU$^2$ Wins: Discovering Efficient Activation Functions for Sparse LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T04:21:30.142851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:21:30.142851Z digest=sha256:5c20182ed1bab7a94351632da35601a8c046747714bf0805c519c3072838dfb6