Pith. sign in

REVIEW 6 cited by

DoctorGLM: Fine-tuning your Chinese Doctor is not a Herculean Task

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.01097 v2 pith:O3EGJ65F submitted 2023-04-03 cs.CL

classification cs.CL
keywords doctorglmmedicalbeenchatgptchinesellmsmodelsa100
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent progress of large language models (LLMs), including ChatGPT and GPT-4, in comprehending and responding to human instructions has been remarkable. Nevertheless, these models typically perform better in English and have not been explicitly trained for the medical domain, resulting in suboptimal precision in diagnoses, drug recommendations, and other medical advice. Additionally, training and deploying a dialogue model is still believed to be impossible for hospitals, hindering the promotion of LLMs. To tackle these challenges, we have collected databases of medical dialogues in Chinese with ChatGPT's help and adopted several techniques to train an easy-deploy LLM. Remarkably, we were able to fine-tune the ChatGLM-6B on a single A100 80G in 13 hours, which means having a healthcare-purpose LLM can be very affordable. DoctorGLM is currently an early-stage engineering attempt and contain various mistakes. We are sharing it with the broader community to invite feedback and suggestions to improve its healthcare-focused capabilities: https://github.com/xionghonglin/DoctorGLM.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SCOPE and SCION: A Benchmark and an Auditable Reference Pipeline for Schema Induction and Fusion from Text

    cs.AI 2026-05 conditional novelty 6.0 of 10

    A 24-dataset benchmark for inducing schema graphs from raw text, plus an auditable LLM-based pipeline that reports the highest scores on the benchmark's four schema-similarity metrics.

  2. V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis

    cs.CE 2025-06 conditional novelty 6.0 of 10

    V2T-CoT combines visual region grounding with LLM-generated text rationale training to improve medical visual question answering accuracy and interpretability on four benchmarks.

  3. DeepForm: Reasoning Large Language Model for Communication System Formulation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    DeepForm, a 7B LLM fine-tuned on the new CSFRC dataset, reports the highest accuracy on a communication system formulation test, surpassing larger models such as DeepSeek R1.

  4. DPF-CM: A Data Processing Framework with Privacy-Preserving Vector Databases for Chinese Medical LLMs Training and Deployment

    cs.LG 2025-09 reject novelty 5.0 of 10

    A data-processing and privacy-preserving deployment framework claims state-of-the-art Chinese medical LLM accuracy and a 27% reduction in training-data leakage.

  5. DiaLLMs: EHR Enhanced Clinical Conversational System for Clinical Test Recommendation and Diagnosis Prediction

    cs.AI 2025-06 conditional novelty 5.0 of 10

    DiaLLM is an EHR-grounded conversational system that translates clinical codes and test results into text and uses PPO with rejection sampling to recommend lab tests and predict diagnoses, reporting large gains over b...

  6. Balancing Knowledge Delivery and Emotional Comfort in Healthcare Conversational Systems

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Fine-tuning a 1B medical chatbot on LLM-rewritten emotional dialogues improves its emotion scores with only small changes in n-gram overlap with the original medical responses.

Pith tools