Pith. sign in

REVIEW 1 cited by

JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.13317 v1 pith:EJJ63KZS submitted 2024-09-20 cs.CL

classification cs.CL
keywords japanesebiomedicalllmsbenchmarkdatasetstasksacrossbetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent developments in Japanese large language models (LLMs) primarily focus on general domains, with fewer advancements in Japanese biomedical LLMs. One obstacle is the absence of a comprehensive, large-scale benchmark for comparison. Furthermore, the resources for evaluating Japanese biomedical LLMs are insufficient. To advance this field, we propose a new benchmark including eight LLMs across four categories and 20 Japanese biomedical datasets across five tasks. Experimental results indicate that: (1) LLMs with a better understanding of Japanese and richer biomedical knowledge achieve better performance in Japanese biomedical tasks, (2) LLMs that are not mainly designed for Japanese biomedical domains can still perform unexpectedly well, and (3) there is still much room for improving the existing LLMs in certain Japanese biomedical tasks. Moreover, we offer insights that could further enhance development in this field. Our evaluation tools tailored to our benchmark as well as the datasets are publicly available in https://huggingface.co/datasets/Coldog2333/JMedBench to facilitate future research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Japanese Language Model and Three New Evaluation Benchmarks for Pharmaceutical NLP

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A continually pretrained 7B Japanese pharmaceutical LLM outperforms open medical models on new Japanese pharma benchmarks, while all models, including GPT-4o, fail at cross-sentence consistency checks.

Pith tools