Pith. sign in

Canonical reference

Comas: Co-evolving multi-agent systems via interaction rewards

Canonical reference. 100% of citing Pith papers cite this work as background.

10 Pith papers citing it
Background 100% of classified citations

citation-role summary

background 5

citation-polarity summary

years

2026 10

roles

background 5

polarities

background 5

representative citing papers

STRIDE: Learnable Stepwise Language Feedback for LLM Reasoning

cs.LG · 2026-05-13 · unverdicted · novelty 6.0

STRIDE co-trains generator and verifier on outcome rewards alone to deliver learnable stepwise language feedback that redirects LLM reasoning trajectories and outperforms scalar-reward baselines.

AIPO: Learning to Reason from Active Interaction

cs.CL · 2026-05-08 · unverdicted · novelty 6.0 · 2 refs

AIPO adds active multi-agent consultation (Verify, Knowledge, Reasoning agents) plus custom importance sampling to RLVR training so LLMs expand their reasoning boundary and then operate without the agents.

citing papers explorer

Showing 10 of 10 citing papers.