Pith. sign in

Attributing mode collapse in the fine-tuning of large language models

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Outcome-based Exploration for LLM Reasoning

cs.LG · 2025-09-08 · conditional · novelty 6.0

Outcome-based exploration bonuses (UCB-Con and Batch) improve pass@1 and pass@32 for LLM math reasoning while slowing diversity collapse, supported by a bandit model with a strong generalization assumption.

citing papers explorer

Showing 1 of 1 citing paper.

  • Outcome-based Exploration for LLM Reasoning cs.LG · 2025-09-08 · conditional · none · ref 2022

    Outcome-based exploration bonuses (UCB-Con and Batch) improve pass@1 and pass@32 for LLM math reasoning while slowing diversity collapse, supported by a bandit model with a strong generalization assumption.