REVIEW 40 cited by
The Forward-Forward Algorithm: Some Preliminary Investigations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The aim of this paper is to introduce a new learning procedure for neural networks and to demonstrate that it works well enough on a few small problems to be worth further investigation. The Forward-Forward algorithm replaces the forward and backward passes of backpropagation by two forward passes, one with positive (i.e. real) data and the other with negative data which could be generated by the network itself. Each layer has its own objective function which is simply to have high goodness for positive data and low goodness for negative data. The sum of the squared activities in a layer can be used as the goodness but there are many other possibilities, including minus the sum of the squared activities. If the positive and negative passes could be separated in time, the negative passes could be done offline, which would make the learning much simpler in the positive pass and allow video to be pipelined through the network without ever storing activities or stopping to propagate derivatives.
Forward citations
Cited by 40 Pith papers
-
Learning in Deep Networks under Dale's Constraint
An on-off two-channel network with fixed-sign synapses and local Hebbian learning is claimed to recover backpropagation exactly under symmetric weights and to beat comparable vanilla networks on Tiny ImageNet.
-
Gauge-Fixing the Forward-Forward Objective: A Whitened Goodness Derived from a Likelihood-Ratio Account
Squared Forward-Forward goodness is the likelihood-ratio statistic for zero-mean populations differing in scale; anisotropic and heavy-tailed cases yield Mahalanobis and saturating (divisive-normalization) forms.
-
Forget, Anticipate and Adapt: Test Time Training for Long Videos
FFN performs TTT on multi-hour videos by restricting updates to three frames and using a surprise metric for adaptive window sizing, plus a new EpicTours dataset.
-
Can Local Learning Match Self-Supervised Backpropagation?
With orthonormal linear networks and optimized layer-wise projections, local-SSL updates equal global BP-SSL updates; adding top-down and 2D-spatial structure to CLAPP then nearly matches BP-SSL on CIFAR-10, STL-10, a...
-
Programmable Photonic Unitary Processor Enables Parametrized Differentiable Long-Haul Spatial Division Multiplexed Transmission
A programmable photonic unitary processor, optimized through a differentiable fiber model, is shown in a 1300-km three-mode fiber experiment to cut modal dispersion by more than half.
-
From Local Learning to Global Prediction Through Layered Surprise Cascades
An inverted Forward-Forward rule makes layered networks cancel expected activity and amplify surprise, producing brain-like bottom-up cascades.
-
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
LoCA tunes adapters with one calibrated map per layer and closed-form ridge solves, avoiding repeated backpropagation after calibration and roughly matching LoRA quality on tested tasks.
-
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
ECL unifies few-shot, continual, zero-shot, and in-context learning on a single TCN-based edge chip, with first hardware baselines on several tasks.
-
Breaking the Curse with BAND: Nonparametric Distribution Estimation in High Dimensions
Sparse Bayesian-network factorization plus sparsity-aware regression yields polynomial TV rates for high-dimensional mixed-type distribution estimation, beating classical histogram rates under sparsity.
-
Reconstructing Backpropagation from Forward Fluctuations in Noise-modulated Neural Networks
In a noise-modulated neural network, backpropagation's transposed-weight multiplications can be replaced by covariance estimates from forward fluctuations, matching backpropagation accuracy on small tasks.
-
Conditioned Direct Feedback Alignment via Activity and Error Geometry
Conditioning the activity and error factors of the direct feedback alignment update with damped inverse second moments improves DFA accuracy in nuisance-dominated regimes and clean confirmations.
-
Backpropagation-Free Trunk Training via the Split Forward Gradients
Split-FG splits a network into an exactly trained head and a forward-gradient-estimated trunk, reducing variance and reaching 387 perplexity on WikiText-103 with a 16M GPT-2-style model.
-
Local Synaptic Rules Can Implement a SIGReg Gradient Without Backpropagation
STDP+ and homeostatic plasticity can compute the exact gradient of a SIGReg-like loss using only local signals, verified in a clustering task and on MNIST.
-
Shunting Inhibition and Dendritic Branching Shape Local Credit Assignment
Exact dendritic gradients factor into local eligibility times path-transported compartment error, and shunting inhibition can improve restricted-feedback local learning by reshaping that error field.
-
Decentralised AI Training and Inference with BlockTrain
BlockTrain partitions models into blocks trained on local objectives, reaching CE 1.359 on WikiText within 0.04 of end-to-end baseline while enabling distributed training and inference over TCP for up to 75B-parameter models.
-
Self-Motivated Growing Neural Network for Adaptive Architecture via Local Structural Plasticity
A gradient-trained control network whose size adjusts online through a local structural plasticity module matches or beats fixed-size MLPs on three control benchmarks.
-
Decoupling Search and Learning in Neural Net Training
Training can be split into evolutionary search over intermediate activations and gradient-based regression to those activations, achieving near-SGD accuracy on three benchmarks.
-
Forward-Only Continual Learning
FoRo achieves strong continual learning accuracy and low forgetting on CIFAR-100, ImageNet-R, and CUB-200 using only forward updates, via CMA-ES prompt tuning and a recursive knowledge encoding matrix.
-
Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis
Reparameterizing layer weights in a learned orthonormal eigenbasis is claimed to improve ImageNet classification, cross-modal retrieval, and enable a faster backpropagation-free variant that surpasses standard backpro...
-
FFGAF-SNN: The Forward-Forward Based Gradient Approximation Free Training Framework for Spiking Neural Networks
A Forward-Forward training framework that freezes spiking layers as black-box encoders and allocates channels by inter-class difficulty achieves 99.58% on MNIST, 92.13% on Fashion-MNIST, and 75.64% on CIFAR-10, the be...
-
Can Biologically Plausible Temporal Credit Assignment Rules Match BPTT for Neural Similarity? E-prop as an Example
At matched task accuracy, e-prop trained RNNs reach neural data similarity comparable to BPTT trained RNNs on Mante 2013 and Sussillo 2015 datasets, with initialization and architecture influencing similarity more tha...
-
Backpropagation-Free Metropolis-Adjusted Langevin Algorithm
Forward-mode automatic differentiation can replace backpropagation inside the Metropolis-Adjusted Langevin Algorithm, yielding four MCMC samplers that are sometimes faster and lower-memory than standard MALA.
-
Forward Target Propagation: A Forward-Only Approach to Global Error Credit Assignment via Local Losses
FTP trains neural networks using only forward passes, propagating random-projection target signals through the network, and achieves near-backpropagation accuracy on small shallow benchmarks.
-
An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models
In CRATE-family transformers, the attention-like MSSA update with skip connection raises the coding rate it was designed to compress, yet layer-averaged SRR still correlates positively with the generalization gap (tau...
-
Train Often, Deploy Selectively: Forward-Gated Model Replacement in Crypto Markets
Shadow Before Swap promotes warm-refit challengers only after a paired one-week NLL edge, beating calendar, blind, and continuous-maintenance baselines while cutting deployments 78%.
-
Beyond Backpropagation: Monte Carlo Method Can Train Deep Neural Networks
Single-parameter Monte Carlo mutation-selection trains deep networks and a simple Transformer on MNIST and Tiny Shakespeare without backpropagation.
-
Steering the Noise: Turning Random Perturbations into Effective Descent for Memory-Efficient LLM Fine-Tuning
Selecting or combining the lowest-loss random perturbations before each update makes zeroth-order LLM fine-tuning converge faster, reportedly beating gradient-based fine-tuning on 9 of 11 tasks at a fraction of the memory.
-
Reshaping the Forward-Forward Algorithm with a Similarity-Based Objective
FAUST replaces the Forward-Forward goodness score with triplet/tuplet similarity losses and achieves near-backpropagation accuracy on MNIST, Fashion-MNIST, and CIFAR-10 with single-pass inference.
-
FF-INT8: Efficient Forward-Forward DNN Training on Edge Devices with INT8 Precision
INT8 quantized training using the Forward-Forward algorithm with a look-ahead loss achieves competitive accuracy and modest resource savings on an edge device.
-
Taxonomic Networks: A Representation for Neuro-Symbolic Pairing
A Cobweb-style symbolic concept learner and a neural soft decision tree are presented as interchangeable 'neuro-symbolic pairs' over taxonomic networks, with complementary data and compute tradeoffs.
-
Scaling of hardware-compatible perturbative training algorithms
Training time to a fixed accuracy for perturbative gradient methods grows far slower than linearly with network size, challenging a long-standing scaling objection.
-
Dendritic Localized Learning: Toward Biologically Plausible Algorithm
DLL trains MLPs, CNNs, and RNNs with local errors and trainable asymmetric feedback, achieving the best accuracy among algorithms satisfying the paper's three biological plausibility criteria.
-
Hardware-In-The-Loop Training of a 4f Optical Correlator with Logarithmic Complexity Reduction for CNNs
Hardware-in-the-loop training of a 4f optical correlator with the PEPITA forward-only algorithm reaches MNIST accuracy comparable to backpropagation while removing the need for software gradients and claiming a log-fa...
-
Scalable Forward-Forward Algorithm
SFF trains convolutional networks layer-by-layer with auxiliary class-goodness layers and block-wise local backpropagation, reaching backprop-comparable accuracy on small benchmarks in the reported runs.
-
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
ZeroFlow is a benchmark showing zeroth-order, forward-pass-only optimizers can match backpropagation-based continual learning on several datasets with about five times lower memory, plus three modest enhancements.
-
Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning
FFCL-SAM, a patch-level classifier plus SAM-based refinement, reports AUC 0.8455 and improved margin segmentation on intraoperative breast radiographs, but the test set excludes negative patients.
-
Training neural networks without backpropagation using particles
Training each neuron of a multilayer perceptron with its own particle-swarm search, instead of using backpropagation, reaches accuracy comparable to gradient-based training on two real datasets and several synthetic ones.
-
A Neural Network Training Method Based on Distributed PID Control
A differential-equation-based training rule with distributed PID control is proposed for a symmetric Wuxing neural network and tested on MNIST.
-
Bootstrapping Life-Inspired Machine Intelligence: The Biological Route from Chemistry to Cognition and Creativity
Flexible, scalable intelligence in biology is attributed to five design principles—autonomy, multiscale self-assembly, continuous rebuilding, embodied constraints, and pervasive signaling—offering a roadmap for life-i...
-
Current Opinions on Memristor-Accelerated Machine Learning Hardware
A review of memristor accelerators for machine learning, covering prototype chips, device/circuit/system challenges, and future directions.
Discussion (0). Continue with ORCID to comment.