Pith. sign in

Superrl: Reinforcement learning with supervision to boost language model reasoning.arXiv preprint arXiv:2506.01096, 2025a

6 Pith papers cite this work. Polarity classification is still indexing.

6 Pith papers citing it

citation-role summary

method 1

citation-polarity summary

years

2026 5 2025 1

roles

method 1

polarities

use method 1

representative citing papers

AIR: Adaptive Interleaved Reasoning with Code in MLLMs

cs.CV · 2026-06-22 · unverdicted · novelty 4.0

AIR applies RL with a group-constrained reward and custom data pipeline to enable adaptive code-interleaved reasoning in MLLMs, reporting 6.1 pp average benchmark gains.

citing papers explorer

Showing 6 of 6 citing papers.