Distribution-Aware Reward optimizes LLM regression by treating rollouts as empirical predictive distributions and rewarding marginal improvements in CRPS quality rather than point accuracy alone.
arXiv preprint arXiv:2402.14547 , year =
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Fine-tuning DeepSeek-R1-1.5B via LoRA on experimental-theoretical deviations yields over 98% training loss reduction and accuracy gains across seven nuclear observables.
citing papers explorer
-
Distribution-Aware Reward: Reinforcement Learning over Predictive Distributions for LLM Regression
Distribution-Aware Reward optimizes LLM regression by treating rollouts as empirical predictive distributions and rewarding marginal improvements in CRPS quality rather than point accuracy alone.
-
Large language model for unified and accurate description of multidimensional nuclear properties
Fine-tuning DeepSeek-R1-1.5B via LoRA on experimental-theoretical deviations yields over 98% training loss reduction and accuracy gains across seven nuclear observables.