{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:PK2ZJANYHHCR3RE7O5RS2BGGT2","short_pith_number":"pith:PK2ZJANY","schema_version":"1.0","canonical_sha256":"7ab59481b839c51dc49f77632d04c69e960210edc583fb045da40aeeb7f1c3ba","source":{"kind":"arxiv","id":"2509.02208","version":1},"attestation_state":"computed","paper":{"title":"Baichuan-M2: Scaling Medical Capability with Large Verifier System","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Baichuan-M2 Team: Chengfeng Dou, Chenzheng Zhu, Chong Liu, Da Pan, Fan Yang, Fei Deng, Fei Li, Guangwei Ai, Guosheng Dong, Hengfu Cui, Hongda Zhang, Jinyang Tai, Jixiang Hong, Jiyuan Jia, Kai Lu, Linzhuang Sun, Mingyang Chen, Peidong Guo, Qiang Ju, Qian Ma, Rihui Xin, Shihui Yang, Shuai Wang, Shunya Dang, Shusen Zhang, Tianpeng Li, Xiangrong Zeng, Xiaochuan Wang, Yichuan Mo, Yijie Zhou, Zheng Liang, Zhishou Zhang, Zuyi Zhu","submitted_at":"2025-09-02T11:23:35Z","abstract_excerpt":"As large language models (LLMs) advance in conversational and reasoning capabilities, their practical application in healthcare has become a critical research focus. However, there is a notable gap between the performance of medical LLMs on static benchmarks such as USMLE and their utility in real-world clinical decision-making. This discrepancy arises because traditional exams fail to capture the dynamic, interactive nature of medical consultations. To address this challenge, we introduce a novel dynamic verification framework that moves beyond static answer verifier, establishing a large-sca"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2509.02208","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.LG","submitted_at":"2025-09-02T11:23:35Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a61d8d97ee5fabde2a7dc13c781c42b241f56f6f00ae4f6e1726de61d06cdae3","abstract_canon_sha256":"7a09531ae31dcd4ef26d911d99719f88b15bb0c69a1995851836fbf5d4e02ea8"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:03:34.056339Z","signature_b64":"r4syxP24XaWjWUQtRYz06hVLuXfYZE8dClNF15ahFtusj2NO2FZQW0qHAykys/SkAJBD/KNg8HO0wAiRaYF2AQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"7ab59481b839c51dc49f77632d04c69e960210edc583fb045da40aeeb7f1c3ba","last_reissued_at":"2026-07-05T12:03:34.055831Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:03:34.055831Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Baichuan-M2: Scaling Medical Capability with Large Verifier System","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.LG","authors_text":"Baichuan-M2 Team: Chengfeng Dou, Chenzheng Zhu, Chong Liu, Da Pan, Fan Yang, Fei Deng, Fei Li, Guangwei Ai, Guosheng Dong, Hengfu Cui, Hongda Zhang, Jinyang Tai, Jixiang Hong, Jiyuan Jia, Kai Lu, Linzhuang Sun, Mingyang Chen, Peidong Guo, Qiang Ju, Qian Ma, Rihui Xin, Shihui Yang, Shuai Wang, Shunya Dang, Shusen Zhang, Tianpeng Li, Xiangrong Zeng, Xiaochuan Wang, Yichuan Mo, Yijie Zhou, Zheng Liang, Zhishou Zhang, Zuyi Zhu","submitted_at":"2025-09-02T11:23:35Z","abstract_excerpt":"As large language models (LLMs) advance in conversational and reasoning capabilities, their practical application in healthcare has become a critical research focus. However, there is a notable gap between the performance of medical LLMs on static benchmarks such as USMLE and their utility in real-world clinical decision-making. This discrepancy arises because traditional exams fail to capture the dynamic, interactive nature of medical consultations. To address this challenge, we introduce a novel dynamic verification framework that moves beyond static answer verifier, establishing a large-sca"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2509.02208","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2509.02208/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2509.02208","created_at":"2026-07-05T12:03:34.055894+00:00"},{"alias_kind":"arxiv_version","alias_value":"2509.02208v1","created_at":"2026-07-05T12:03:34.055894+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2509.02208","created_at":"2026-07-05T12:03:34.055894+00:00"},{"alias_kind":"pith_short_12","alias_value":"PK2ZJANYHHCR","created_at":"2026-07-05T12:03:34.055894+00:00"},{"alias_kind":"pith_short_16","alias_value":"PK2ZJANYHHCR3RE7","created_at":"2026-07-05T12:03:34.055894+00:00"},{"alias_kind":"pith_short_8","alias_value":"PK2ZJANY","created_at":"2026-07-05T12:03:34.055894+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":10,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08257","citing_title":"MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11740","citing_title":"UniReason-Med: A Shared Grounded Reasoning Interface for 2D-to-3D Transfer in Medical VQA","ref_index":274,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09365","citing_title":"Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2606.03157","citing_title":"ClinicalMC: A Benchmark for Multi-Course Clinical Decision-Making with Large Language Models","ref_index":52,"is_internal_anchor":false},{"citing_arxiv_id":"2606.29928","citing_title":"Latent-CURE for Breast Cancer Diagnosis","ref_index":22,"is_internal_anchor":false},{"citing_arxiv_id":"2605.27860","citing_title":"C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08982","citing_title":"Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care","ref_index":2,"is_internal_anchor":false},{"citing_arxiv_id":"2606.11675","citing_title":"Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning","ref_index":53,"is_internal_anchor":false},{"citing_arxiv_id":"2509.08827","citing_title":"A Survey of Reinforcement Learning for Large Reasoning Models","ref_index":221,"is_internal_anchor":false},{"citing_arxiv_id":"2604.08559","citing_title":"Medical Reasoning with Large Language Models: A Survey and MR-Bench","ref_index":43,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2","json":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2.json","graph_json":"https://pith.science/api/pith-number/PK2ZJANYHHCR3RE7O5RS2BGGT2/graph.json","events_json":"https://pith.science/api/pith-number/PK2ZJANYHHCR3RE7O5RS2BGGT2/events.json","paper":"https://pith.science/paper/PK2ZJANY"},"agent_actions":{"view_html":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2","download_json":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2.json","view_paper":"https://pith.science/paper/PK2ZJANY","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2509.02208&json=true","fetch_graph":"https://pith.science/api/pith-number/PK2ZJANYHHCR3RE7O5RS2BGGT2/graph.json","fetch_events":"https://pith.science/api/pith-number/PK2ZJANYHHCR3RE7O5RS2BGGT2/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2/action/timestamp_anchor","attest_storage":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2/action/storage_attestation","attest_author":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2/action/author_attestation","sign_citation":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2/action/citation_signature","submit_replication":"https://pith.science/pith/PK2ZJANYHHCR3RE7O5RS2BGGT2/action/replication_record"}},"created_at":"2026-07-05T12:03:34.055894+00:00","updated_at":"2026-07-05T12:03:34.055894+00:00"}