Combining continued pretraining with reasoning preference optimization yields a 72B Japanese medical model that keeps 0.868 accuracy on IgakuQA with and without explanation prompting, while a model without RPO drops to 0.834.
PLaMo-100B: A Ground-Up Language Model Designed for Japanese Proficiency
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We introduce PLaMo-100B, a large-scale language model designed for Japanese proficiency. The model was trained from scratch using 2 trillion tokens, with architecture such as QK Normalization and Z-Loss to ensure training stability during the training process. Post-training techniques, including Supervised Fine-Tuning and Direct Preference Optimization, were applied to refine the model's performance. Benchmark evaluations suggest that PLaMo-100B performs well, particularly in Japanese-specific tasks, achieving results that are competitive with frontier models like GPT-4. The base model is available at https://huggingface.co/pfnet/plamo-100b.
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Stabilizing Reasoning in Medical LLMs with Continued Pretraining and Reasoning Preference Optimization
Combining continued pretraining with reasoning preference optimization yields a 72B Japanese medical model that keeps 0.868 accuracy on IgakuQA with and without explanation prompting, while a model without RPO drops to 0.834.