SFT and RL cannot be decoupled in LLM post-training because each step increases the loss or lowers the reward of the prior step under KL and PL analyses.
Beyond scaling laws: Understanding transformer performance with associative memory
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
APRTrack applies hierarchical adversarial perturbations at modality and spatial levels plus footprint-calibrated Hopfield retrieval to improve robustness of RGB-Event tracking under occlusion and modal failure.
Proposes a semantic information theory for LLMs that substitutes the token for the bit as the atomic carrier of meaning, recasts the Transformer as an energy-based model, and derives directed rate-distortion and rate-reward functions using Massey's directed information.
citing papers explorer
-
On the Non-decoupling of Supervised Fine-tuning and Reinforcement Learning in Post-training
SFT and RL cannot be decoupled in LLM post-training because each step increases the loss or lowers the reward of the prior step under KL and PL analyses.
-
Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking
APRTrack applies hierarchical adversarial perturbations at modality and spatial levels plus footprint-calibrated Hopfield retrieval to improve robustness of RGB-Event tracking under occlusion and modal failure.
-
Forget BIT, It is All about TOKEN: Towards Semantic Information Theory for LLMs
Proposes a semantic information theory for LLMs that substitutes the token for the bit as the atomic carrier of meaning, recasts the Transformer as an energy-based model, and derives directed rate-distortion and rate-reward functions using Massey's directed information.