Pith. sign in

The hallucination tax of rein- forcement finetuning.arXiv preprint arXiv:2505.13988

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 3 cs.CL 1

years

2026 4

roles

background 1

polarities

background 1

representative citing papers

Scaling Participation in Modular AI Systems

cs.AI · 2026-06-05 · unverdicted · novelty 6.0

Modular AI systems assembled from contributed small models outperform monolithic LLMs by up to 15.4% on 15 tasks including reasoning and factuality while showing emergent problem-solving and benefits from contributor diversity.

MoCo: A One-Stop Shop for Model Collaboration Research

cs.CL · 2026-01-29 · accept · novelty 6.0

MoCo supplies a unified library of 26 collaboration strategies and benchmarks demonstrating average outperformance over single models in 61 percent of (model, data) pairs.

citing papers explorer

Showing 4 of 4 citing papers.

  • Scaling Participation in Modular AI Systems cs.AI · 2026-06-05 · unverdicted · none · ref 81

    Modular AI systems assembled from contributed small models outperform monolithic LLMs by up to 15.4% on 15 tasks including reasoning and factuality while showing emergent problem-solving and benefits from contributor diversity.

  • MoCo: A One-Stop Shop for Model Collaboration Research cs.CL · 2026-01-29 · accept · none · ref 21

    MoCo supplies a unified library of 26 collaboration strategies and benchmarks demonstrating average outperformance over single models in 61 percent of (model, data) pairs.

  • Bridging the Detection-to-Abstention Gap in Reasoning Models under Insufficient Information cs.AI · 2026-05-27 · unverdicted · none · ref 42

    JTS trains reasoning models via supervised warm-up and missing-premise RL to make an explicit answerability commitment that triggers early termination on unanswerable inputs, raising Abstention@Detection near saturation.

  • Rethinking Agentic Reinforcement Learning In Large Language Models cs.AI · 2026-04-30 · unverdicted · none · ref 79 · 3 links

    The paper reviews conceptual foundations, methodological innovations, effective designs, critical challenges, and future directions for LLM-based Agentic Reinforcement Learning.