HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.
Towards next-generation medical agent: How o1 is reshaping decision-making in medical scenarios
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CL 2verdicts
UNVERDICTED 2representative citing papers
HiMed releases a Hindi medical reasoning corpus and benchmark and shows that training an 8B LLM with decaying scaffolding reward improves Hindi performance and narrows the English-Hindi accuracy gap.
citing papers explorer
-
HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
HuatuoGPT-o1 achieves superior medical complex reasoning by using a verifier to curate reasoning trajectories for fine-tuning and then applying RL with verifier-based rewards.
-
HiMed: Incentivizing Hindi Reasoning in Medical LLMs
HiMed releases a Hindi medical reasoning corpus and benchmark and shows that training an 8B LLM with decaying scaffolding reward improves Hindi performance and narrows the English-Hindi accuracy gap.