Pith. sign in

Scaling of search and learning: A roadmap to reproduce o1 from reinforcement learning perspective

7 Pith papers cite this work, alongside 4 external citations. Polarity classification is still indexing.

7 Pith papers citing it
4 external citations · external index

citation-role summary

background 2

citation-polarity summary

fields

cs.CL 5 cs.AI 2

roles

background 2

polarities

background 2

representative citing papers

Search-o1: Agentic Search-Enhanced Large Reasoning Models

cs.AI · 2025-01-09 · unverdicted · novelty 6.0

Search-o1 integrates agentic retrieval-augmented generation and a Reason-in-Documents module into large reasoning models to dynamically supply missing knowledge and improve performance on complex science, math, coding, and QA tasks.

HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs

cs.CL · 2024-12-25 · unverdicted · novelty 6.0

HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.

citing papers explorer

Showing 7 of 7 citing papers.