Pith. sign in

REVIEW 3 cited by

Feed-Forward Networks with Attention Can Solve Some Long-Term Memory Problems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1512.08756 v5 pith:ASI366PQ submitted 2015-12-29 cs.LG cs.NE

classification cs.LGcs.NE
keywords attentionfeed-forwardlong-termmemorymodelnetworksproblemssolve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a simplified model of attention which is applicable to feed-forward neural networks and demonstrate that the resulting model can solve the synthetic "addition" and "multiplication" long-term memory problems for sequence lengths which are both longer and more widely varying than the best published results for these tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Eigenvalues as a Metric for Memory Dynamics in Sequence Models

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Eigenvalue spectra of attention and SSM dynamics show consistent signatures of memory retention and selective forgetting that align with task requirements.

  2. An Attention Mechanism for Musical Instrument Recognition

    cs.IR 2019-07 unverdicted novelty 5.0 of 10

    An attention model improves multi-label instrument recognition accuracy on the weakly labeled OpenMIC dataset compared to baseline, RNN, and fully connected networks across 20 instruments.

  3. Attention based Bidirectional GRU hybrid model for inappropriate content detection in Urdu language

    cs.CL 2025-01 conditional novelty 4.0 of 10

    An attention-based BiGRU model achieves 84% accuracy on a new combined Urdu inappropriate-content dataset, outperforming four baseline deep learning models without pre-trained word vectors.

Pith tools