Exact Flow Linear Attention derives a closed-form exact update for delta-rule linear attention from continuous-time dynamics, removing Euler discretization error while preserving linear complexity and structure.
Mega: Moving average equipped gated attention
7 Pith papers cite this work, alongside 36 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Self-pretraining improves Transformer sequence classification by enabling learning of proximity-biased attention from positional encodings that label supervision alone cannot easily acquire from random starts.
Pre-training ten DNN architectures on knowledge-driven synthetic ECGs generated via Gaussian PQRST wave composition improves classification of AF, AFLT, PVC, and WPW, with largest gain of 33.2% for AFLT and stronger benefits on smaller real datasets.
LLMs memorize citations hierarchically: titles and first authors are recalled at lower redundancy levels than venues or years, with accuracy scaling log-linearly and saturating near verbatim reproduction above roughly 1200 citations.
DemaFormer pairs energy-based modeling with a damped-EMA Transformer to localize video moments matching language queries and reports gains over baselines on four datasets.
citing papers explorer
-
Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
Exact Flow Linear Attention derives a closed-form exact update for delta-rule linear attention from continuous-time dynamics, removing Euler discretization error while preserving linear complexity and structure.
-
Towards Understanding Self-Pretraining for Sequence Classification
Self-pretraining improves Transformer sequence classification by enabling learning of proximity-biased attention from positional encodings that label supervision alone cannot easily acquire from random starts.
-
Boosting ECG Classification Performance by Pre-training with Synthesized Data
Pre-training ten DNN architectures on knowledge-driven synthetic ECGs generated via Gaussian PQRST wave composition improves classification of AF, AFLT, PVC, and WPW, with largest gain of 33.2% for AFLT and stronger benefits on smaller real datasets.
-
Hierarchical Memorization in Large Language Models: Evidence from Citation Generation
LLMs memorize citations hierarchically: titles and first authors are recalled at lower redundancy levels than venues or years, with accuracy scaling log-linearly and saturating near verbatim reproduction above roughly 1200 citations.
-
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
DemaFormer pairs energy-based modeling with a damped-EMA Transformer to localize video moments matching language queries and reports gains over baselines on four datasets.
- The Impossibility Triangle of Long-Context Modeling
- Advancing Intelligent Sequence Modeling: Evolution, Trade-offs, and Applications of State-Space Architectures from S4 to Mamba