Adding temporal memory via LIF, precision-weighted gating, and anticipatory prediction to MoE routers recovers effective expert selection at distribution transitions, with ablation confirming a super-additive beta-ant interaction.
Predictive coding: Towards a future of deep learning beyond backpropagation?
4 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 4roles
background 1polarities
background 1representative citing papers
Predictive coding is recast as deep hierarchical Gaussian filters to restore precision-weighted message passing, yielding closed-form inference and online precision learning that matches backpropagation speed on FashionMNIST while outperforming on online and concept-drift tasks.
Coupled entropy maximized by coupled stretched exponentials is the unique universal entropy meeting scale-specific uncertainty measurement and Hanel-Thurner extensivity requirements.
The book presents principles from optimization and information theory to explain deep network architectures and enable new interpretable models.
citing papers explorer
-
Affinity Is Not Enough: Recovering the Free Energy Principle in Mixture-of-Experts
Adding temporal memory via LIF, precision-weighted gating, and anticipatory prediction to MoE routers recovers effective expert selection at distribution transitions, with ablation confirming a super-additive beta-ant interaction.
-
Closed-form predictive coding via hierarchical Gaussian filters
Predictive coding is recast as deep hierarchical Gaussian filters to restore precision-weighted message passing, yielding closed-form inference and online precision learning that matches backpropagation speed on FashionMNIST while outperforming on online and concept-drift tasks.
-
The unique, universal entropy for complex systems
Coupled entropy maximized by coupled stretched exponentials is the unique universal entropy meeting scale-specific uncertainty measurement and Hanel-Thurner extensivity requirements.
-
Principles and Practice of Deep Representation Learning: or a Mathematical Theory of Memory
The book presents principles from optimization and information theory to explain deep network architectures and enable new interpretable models.