Introduces TA-MDP and proves GRPO convergence at O(1/sqrt(T)), a reward decomposition bound, and PAC-Bayes generalization for tool-augmented LVLM policies.
arXiv preprint arXiv:2410.19732 , year=
2 Pith papers cite this work. Polarity classification is still indexing.
years
2026 2verdicts
UNVERDICTED 2representative citing papers
HierBias introduces a context-conditioned hierarchical architecture with theoretical bounds showing context reduces Bayes error and multi-task learning for bias detection and type classification, reporting improved F1 and MCC on BABE and BASIL datasets.
citing papers explorer
-
Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization
Introduces TA-MDP and proves GRPO convergence at O(1/sqrt(T)), a reward decomposition bound, and PAC-Bayes generalization for tool-augmented LVLM policies.
-
HierBias: Context-Conditioned Hierarchical Media Bias Detection with Multi-Task Type Classification
HierBias introduces a context-conditioned hierarchical architecture with theoretical bounds showing context reduces Bayes error and multi-task learning for bias detection and type classification, reporting improved F1 and MCC on BABE and BASIL datasets.