Pith. sign in

hub Canonical reference

Deepseek-r1 incentivizes reasoning in llms through reinforcement learning.Nature, 645(8081):633–638

Canonical reference. 83% of citing Pith papers cite this work as background.

22 Pith papers citing it
Background 83% of classified citations

hub tools

citation-role summary

background 5 other 1

citation-polarity summary

years

2026 21 2025 1

polarities

background 5 unclear 1

representative citing papers

How Post-Training Shapes Biological Reasoning Models

cs.LG · 2026-06-15 · unverdicted · novelty 6.0

Post-training stages reshape generalization in biological reasoning models distinctly: CPT aligns with biological language, SFT boosts ID performance but causes OOD to peak early and decline, while RL on strong SFT checkpoints can recover OOD generalization.

Stateful Reasoning via Insight Replay

cs.AI · 2026-05-14 · unverdicted · novelty 6.0 · 2 refs

InsightReplay improves long CoT reasoning by extracting critical insights from the trace and replaying them near the active frontier, delivering +1.65 average accuracy gain across 24 model-benchmark settings.

MindLoom: Composing Thought Modes for Frontier-Level Reasoning Data Synthesis

cs.AI · 2026-05-20 · unverdicted · novelty 5.0

MindLoom synthesizes frontier-level reasoning data by decomposing solutions into thought mode chains, training a retrieval model for mode selection, composing new problems with distribution-aligned sampling, and applying rollout-based difficulty labeling for fine-tuning.

Tabular Foundation Model for Generative Modelling

cs.LG · 2026-05-10 · unverdicted · novelty 5.0

TabFORGE generates high-quality synthetic tabular data by leveraging pretrained causality-aware representations in a two-stage diffusion-decoder architecture that mitigates latent distribution shifts.

citing papers explorer

Showing 22 of 22 citing papers.