A framework to identify and convert foldable layer normalizations to RMSNorm for exact equivalence and faster inference in deep neural networks.
Title resolution pending
13 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 13roles
background 3polarities
background 3representative citing papers
LLMs process negation using both attention-based suppression and constructive representation mechanisms (construction dominant), with late-layer attention shortcuts explaining poor accuracy on negation tasks.
A Deep Set encoder plus normalizing flow model trained on five million CRPropa 3 events recovers UHECR source parameters without bias and classifies primary composition at over 98 percent accuracy.
ResRL decouples shared semantics between positive and negative responses in LLM reinforcement learning via SVD-based projection residuals, outperforming baselines including NSR by up to 9.4% on math reasoning benchmarks.
SDM sparsifies the Gated DeltaNet update rule to enable 1000x larger recurrent memory states at iso-FLOP, improving long-context recall and short-context reasoning over GDN and matching full attention at 8B scale.
The work augments pose-conditioned 3D Gaussian avatars with a residual latent evolved by a transformer decoder that decomposes updates into driving, restoring, and dissipative forces to produce history-dependent, temporally coherent full-body animations.
QED bounds cross-run KL divergence in Boltzmann policies by setting temperature proportional to Q-disagreement and reduces return variance by two orders of magnitude on 18 continuous-control tasks without performance loss.
STELLAR trains up to 500M-parameter multi-modal models on 50M driving scenes and reports empirical scaling trends plus new state-of-the-art results on the Waymo Open Dataset.
Double metric learning learns two embeddings per node to build directed graphs with chain connections, yielding better performance than single metric learning for high-pT particles and accurate edge direction prediction in ATLAS ITk simulations.
The paper casts the standard Transformer block with RoPE as a first-order approximation of a radial–tangential state estimator and introduces a Polar Transformer variant that retains the discarded geometric corrections.
mlr3torch introduces an extensible deep learning framework in R that integrates torch models into the mlr3 ecosystem via graph-based architectures for classification, regression, and multimodal tasks.
CMRU restores gradient flow in BMRU via cumulative state updates with skip-connections through time, yielding better convergence and benchmark performance while retaining quantized persistent memory.
Tuned classic GNNs outperform specialized multi-label node classification methods on four of five benchmarks and reach state-of-the-art in multiple settings.
citing papers explorer
-
Enjoy Your Layer Normalization with the Computational Efficiency of RMSNorm
A framework to identify and convert foldable layer normalizations to RMSNorm for exact equivalence and faster inference in deep neural networks.
-
How Language Models Process Negation
LLMs process negation using both attention-based suppression and constructive representation mechanisms (construction dominant), with late-layer attention shortcuts explaining poor accuracy on negation tasks.
-
Neural Posterior Estimation for UHECR source inference from 3D propagation simulations
A Deep Set encoder plus normalizing flow model trained on five million CRPropa 3 events recovers UHECR source parameters without bias and classifies primary composition at over 98 percent accuracy.
-
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning
ResRL decouples shared semantics between positive and negative responses in LLM reinforcement learning via SVD-based projection residuals, outperforming baselines including NSR by up to 9.4% on math reasoning benchmarks.
-
Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsity
SDM sparsifies the Gated DeltaNet update rule to enable 1000x larger recurrent memory states at iso-FLOP, improving long-context recall and short-context reasoning over GDN and matching full attention at 8B scale.
-
Latent Dynamics for Full Body Avatar Animation
The work augments pose-conditioned 3D Gaussian avatars with a residual latent evolved by a transformer decoder that decomposes updates into driving, restoring, and dissipative forces to produce history-dependent, temporally coherent full-body animations.
-
Behavior-Consistent Deep Reinforcement Learning
QED bounds cross-run KL divergence in Boltzmann policies by setting temperature proportional to Q-disagreement and reduces return variance by two orders of magnitude on 18 continuous-control tasks without performance loss.
-
STELLAR: Scaling 3D Perception Large Models for Autonomous Driving
STELLAR trains up to 500M-parameter multi-modal models on 50M driving scenes and reports empirical scaling trends plus new state-of-the-art results on the Waymo Open Dataset.
-
Double Metric Learning for Building Directed Graphs with Chain Connections for the ATLAS ITk Detector
Double metric learning learns two embeddings per node to build directed graphs with chain connections, yielding better performance than single metric learning for high-pT particles and accurate edge direction prediction in ATLAS ITk simulations.
-
The Transformer as a Polar State Estimator
The paper casts the standard Transformer block with RoPE as a first-order approximation of a radial–tangential state estimator and introduces a Polar Transformer variant that retains the discarded geometric corrections.
-
mlr3torch: A Deep Learning Framework in R based on mlr3 and torch
mlr3torch introduces an extensible deep learning framework in R that integrates torch models into the mlr3 ecosystem via graph-based architectures for classification, regression, and multimodal tasks.
-
Improving the Performance and Learning Stability of Parallelizable RNNs Designed for Ultra-Low Power Applications
CMRU restores gradient flow in BMRU via cumulative state updates with skip-connections through time, yielding better convergence and benchmark performance while retaining quantized persistent memory.
-
Rethinking Multi-Label Node Classification: Do Tuned Classic GNNs Suffice?
Tuned classic GNNs outperform specialized multi-label node classification methods on four of five benchmarks and reach state-of-the-art in multiple settings.