Pith. sign in

REVIEW 2 cited by

Fast Feedforward Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.14711 v2 pith:WOBZIJ54 submitted 2023-08-28 cs.LG cs.AIcs.PF

classification cs.LGcs.AIcs.PF
keywords feedforwardnetworksfastfasterfffsinferencelayeralternative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We break the linear link between the layer size and its inference cost by introducing the fast feedforward (FFF) architecture, a log-time alternative to feedforward networks. We demonstrate that FFFs are up to 220x faster than feedforward networks, up to 6x faster than mixture-of-experts networks, and exhibit better training properties than mixtures of experts thanks to noiseless conditional execution. Pushing FFFs to the limit, we show that they can use as little as 1% of layer neurons for inference in vision transformers while preserving 94.2% of predictive performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Gated amplitude-only FFN interventions improve tool-structured LLM outputs on Qwen models by several points, while direction-changing repairs harm more than they fix.

  2. Position: A Theory of Deep Learning Must Include Compositional Sparsity

    cs.LG 2025-07 conditional novelty 4.0 of 10

    All polynomial-time computable functions are compositionally sparse, and this property is the proposed reason deep networks avoid the curse of dimensionality and achieve practical success.

Pith tools