{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:HFOM2IKPYPXZLR6PYABUHN7IR7","short_pith_number":"pith:HFOM2IKP","schema_version":"1.0","canonical_sha256":"395ccd214fc3ef95c7cfc00343b7e88fd4cb2afc029a60c55fd3951dcf9c0c45","source":{"kind":"arxiv","id":"2506.04078","version":3},"attestation_state":"computed","paper":{"title":"LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Binze Hu, Changhao Jiang, Chenhao Huang, Huayu Sha, Jingqi Tong, Mingxu Chai, Ming Zhang, Qi Zhang, Shichun Liu, Shihan Dou, Tao Gui, Xuanjing Huang, Yuhui Wang, Yujiong Shen, Zelin Li, Zhiheng Xi","submitted_at":"2025-06-04T15:43:14Z","abstract_excerpt":"Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and specialized assessments. However, these benchmarks have limitations in question design (mostly multiple-choice), data sources (often not derived from real clinical scenarios), and evaluation methods (poor assessment of complex reasoning). To address these issues, we present LLMEval-Med, a new benchmark covering five core medical areas, including 2,996 questio"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2506.04078","kind":"arxiv","version":3},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2025-06-04T15:43:14Z","cross_cats_sorted":["cs.AI"],"title_canon_sha256":"a74da155a1278ac41f6affb9be3df976d1897e0740354fcc7634d1ce08b9b4bd","abstract_canon_sha256":"b4b681ea9e0af73a7b8e3a0c788ed5068cebe823538c2696205d7b3b554f8afe"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T12:02:02.366521Z","signature_b64":"SLn+y/tP/kZmfALMYtBvZOZbOKjmTeTaHmYscoTySP4HZU/eOj87LOFOA2g3M85+GiLXHUwT2TMj8QSNTzCbBg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"395ccd214fc3ef95c7cfc00343b7e88fd4cb2afc029a60c55fd3951dcf9c0c45","last_reissued_at":"2026-07-05T12:02:02.365994Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T12:02:02.365994Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Binze Hu, Changhao Jiang, Chenhao Huang, Huayu Sha, Jingqi Tong, Mingxu Chai, Ming Zhang, Qi Zhang, Shichun Liu, Shihan Dou, Tao Gui, Xuanjing Huang, Yuhui Wang, Yujiong Shen, Zelin Li, Zhiheng Xi","submitted_at":"2025-06-04T15:43:14Z","abstract_excerpt":"Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and specialized assessments. However, these benchmarks have limitations in question design (mostly multiple-choice), data sources (often not derived from real clinical scenarios), and evaluation methods (poor assessment of complex reasoning). To address these issues, we present LLMEval-Med, a new benchmark covering five core medical areas, including 2,996 questio"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2506.04078","kind":"arxiv","version":3},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2506.04078/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.04078","created_at":"2026-07-05T12:02:02.366055+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.04078v3","created_at":"2026-07-05T12:02:02.366055+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.04078","created_at":"2026-07-05T12:02:02.366055+00:00"},{"alias_kind":"pith_short_12","alias_value":"HFOM2IKPYPXZ","created_at":"2026-07-05T12:02:02.366055+00:00"},{"alias_kind":"pith_short_16","alias_value":"HFOM2IKPYPXZLR6P","created_at":"2026-07-05T12:02:02.366055+00:00"},{"alias_kind":"pith_short_8","alias_value":"HFOM2IKP","created_at":"2026-07-05T12:02:02.366055+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":4,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2607.08257","citing_title":"MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters","ref_index":34,"is_internal_anchor":true},{"citing_arxiv_id":"2606.28332","citing_title":"When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2605.30637","citing_title":"EHRBench: An Automated and Reliable EHR-based Benchmark for Clinical Decision Making with LLMs","ref_index":123,"is_internal_anchor":false},{"citing_arxiv_id":"2509.22258","citing_title":"Beyond Classification Accuracy: Neural-MedBench and the Need for Deeper Reasoning Benchmarks","ref_index":34,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7","json":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7.json","graph_json":"https://pith.science/api/pith-number/HFOM2IKPYPXZLR6PYABUHN7IR7/graph.json","events_json":"https://pith.science/api/pith-number/HFOM2IKPYPXZLR6PYABUHN7IR7/events.json","paper":"https://pith.science/paper/HFOM2IKP"},"agent_actions":{"view_html":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7","download_json":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7.json","view_paper":"https://pith.science/paper/HFOM2IKP","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.04078&json=true","fetch_graph":"https://pith.science/api/pith-number/HFOM2IKPYPXZLR6PYABUHN7IR7/graph.json","fetch_events":"https://pith.science/api/pith-number/HFOM2IKPYPXZLR6PYABUHN7IR7/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7/action/timestamp_anchor","attest_storage":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7/action/storage_attestation","attest_author":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7/action/author_attestation","sign_citation":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7/action/citation_signature","submit_replication":"https://pith.science/pith/HFOM2IKPYPXZLR6PYABUHN7IR7/action/replication_record"}},"created_at":"2026-07-05T12:02:02.366055+00:00","updated_at":"2026-07-05T12:02:02.366055+00:00"}