Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T00:40:34.440305Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 9 inbound Pith citation observations for arXiv:2508.04325.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T00:40:34.440305Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:57:40.893213Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
57 of 57 outbound references displayed
External citation measurements
3
pith, observed 2026-08-05T02:28:24.338817Z
Observation 1882affb-a6d0-438d-9f31-9e717c4cdad5 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Extending Internet Access Over LoRa for Internet of Things and Critical Applications
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6a12dde6-efe5-4dfd-a678-8026e8bc1c33 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c83eae94-12c4-46ae-9aec-cc52f427a514 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models A Survey on Data Contamination for Large Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 967d7881-6b0a-438b-b28a-027b74ebfe1d · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c475641c-7c69-4ecf-aa9a-02a0b10a969f · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models COGNET-MD, an evaluation framework and dataset for Large Language Model benchmarks in the medical domain
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 58f1b971-5115-4972-88cc-187979c2b949 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc344051-3cd2-4e62-9f3f-9bff7da01e77 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 265a1e35-1630-4cbf-82af-bbea73f97d58 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Ndepartment de- notes the number of medical departments, typi- cally referring to the medical specialties included in the model’s evaluation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 254418af-93aa-4d7c-8828-d326e85dc3c7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46a46edf-669b-4cf1-8613-c7a638dc6533 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models The remaining 72% (38 out of 53) did not conduct in- ternal consistency assessments
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f15c379-6bad-4455-a54f-306ba260ab66 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models These findings highlight a lack of rigorous statistical stan- dards in current benchmark design
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd55380f-431d-4d8f-8da6-218be68a6ebd · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clearly defined evaluation objectives can avoid ambiguity, facili- tating subsequent data collection, task design, and metric selection
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6e578f29-ce02-4261-b9ba-924cee9406f3 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Linking the benchmark to real-world application scenarios en- sures that the results are meaningful
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9c5e1c6-3b52-4773-8332-3a9af8507bd7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Demonstrating the unique- ness demonstrates the necessity and jus- tification of the new benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8612c382-dbd2-4134-ab22-61234d9a245c · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Scoring: – 0: Does not define the target LLM capability
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09c00e7b-7395-4424-b2e6-54d32ef3e37e · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: By clearly defining the medical scope, it helps users better un- derstand the breath and depth of the cov- erage of the benchmark
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 115ae008-ebe8-4610-ac17-0acbdd4448d7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: An effective benchmark should serve the needs of users
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation afa28656-4f21-458a-9a2f-5661ac20abdd · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Due to the professionalism and rigor required in the medical field, the development of a benchmark must involve deep engagement from domain experts
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af582f06-69b6-443b-bfa3-51f07b205f63 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Benchmark content should adhere to recognized, evidence- based medical knowledge sources
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 54b06b0b-bfd1-43e8-9dfc-3a22d65cfb31 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Adherence to medical standards ensures the clinical relevance and consistency, facilitating integration and comparison in reflect real-world medical practice
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a0d267c-a790-4a4e-8775-45d0295c4742 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Evaluation metrics di- rectly shape the interpretation of results
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d837b2a-c9ec-41d3-9b41-993b9ea59a8c · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: In high-stakes medical do- main, going beyond correctness is vi- tal
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90ff61c2-0a41-4b02-ae99-c208046edd83 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a4f054ea-2641-43f2-ad86-1ad89a34ad89 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear and traceable data sources are critical to ensure trans- parency and ethical data usage, which is especially important when sensitive data is involved
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2cd13a52-7c2e-44bf-8c0e-b6a43ef6218b · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Collecting data from unre- liable sources may lead to invalid results for medical applications
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 26248f49-8d28-4ac9-83b7-388909d07d43 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models For synthetically generated data, the con- struction process and verification for its authenticity (e.g., expert review) should also be described
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da50aa52-0560-44e8-8d6f-a1ac0c9946f0 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A benchmark that lacks representativeness may lead to bias in evaluation, reducing the clinical rele- vance, generalizability and fairness of the results
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5994d96b-e7f1-4614-928e-423073eec76f · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Ensuring the dataset cov- ers a variety helps comprehensively eval- uate the model’s generalization ability, reducing bias in the evaluation results
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e42c3370-851a-4a00-b9b6-417e3ba13e45 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that the final dataset is well-structured, enhancing re- liability
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa223e34-7247-4c85-95a6-cfdd5525fc8c · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Methods of de-identification should be described and compliance with relevant regulations (e.g., HIPAA) should be clearly stated
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 50fc5e83-57f2-4993-9827-6357b3053bb2 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A clear and consistent for- mat is essential for standardized evalua- tion
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4505880c-2f99-471d-812f-24e410f96e33 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A dataset construction pro- cess without review mechanism is prone to errors
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bea2e980-b546-4663-b09a-922cebbd5a73 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear reference answers or scoring guidelines ensures transpar- ent and accurate evaluation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a75d24a2-b9f0-46d3-a2a4-eea584efc0c5 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Data contamination may lead to inflated performance, which only reflect memorization instead of medical capability from the models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b31b3c6d-4ae9-4238-bf66-8ce3282ff151 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that users can use the benchmark conveniently, promot- ing benchmark adoption and ensuring fair, transparent, and consistent evalua- tion
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e5e6522-2d9a-4b6f-928b-c0510cbf89a5 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90d60789-305f-4f8a-8114-8bcee9f2df84 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Providing different per- formance baselines allows comparison against the model’s performance, en- abling a deeper understanding and better interpretability
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0b27be8a-1277-44db-9437-4516e2592798 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: In medical domain, un- derstanding the model’s decision-making process is just as important as the final answer
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c4670c0-ee98-4236-a665-adff6ff18ef4 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Robustness testing en- sures model’s output is consistent and reliable under different conditions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e559e94-892d-4108-9972-9c703ea4c069 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a87bf755-a425-4f44-bc89-28e74f1ad7f0 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models I don’t know
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 75f581da-0c51-41ed-9e84-e3cb2fb9927b · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: It ensures that different types of models can be tested under the same interface
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6938e0c3-ce32-4bac-9b7d-17a2803e2ef7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Sufficient coverage of the core medical competencies the bench- mark aims to measure is the prerequisite for establishing content validity
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe874795-7e6f-4106-919b-ae7a965300c3 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Ensuring that the evalu- ation task closely mirrors the targeted clinical practice in real-world scenarios enhance the relevance of the benchmark
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 237b8ce0-284e-4b9a-86ec-b3ae24123eb9 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: An effective benchmark should be capable of differentiating and distinguishing models of varying capabil- ities
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 448593d2-abb6-4985-9034-af5d5e4f029b · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07a93eb3-cf09-4d01-bfea-f096a83fd057 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Scoring: – 0: Does not mention or conduct any internal consistency measurement
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 57553101-f584-42f9-a6cb-227bb695d59f · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1f8582eb-a03b-49b5-9826-2d23b583ec2a · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: A complete and clear doc- umentation help users understand and use the benchmark properly, enhancing the usability, reproducibility and trans- parency
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e086260e-9abc-4c22-bbba-40f112dbed87 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clear evaluation guide- lines help users better understand how model performance is quantified, ensur- ing a shared interpretation and under- standing
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c6186465-86d9-4fe7-8d99-e82efc9804e6 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Disclosing limitations and risks demonstrates scientific rigor and responsibility
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7322dbd1-2315-4f61-981c-143649b8df95 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Going through peer review process means that the design, validity and results of a benchmark has been rig- orous evaluated
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d37a8ac-628e-4a4c-8b71-b39960abd0b8 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models on platforms like GitHub or Hugging Face) along with the applica- ble license
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1a4caca-eddf-4e7d-9e5e-3b183afb2fb7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Proper usage and citation guidelines help maintain academic in- tegrity
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9e4bcdd5-65eb-4968-8231-a82ffecb67bf · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42ffb1e8-6d51-4cbe-beb4-5696a4053aa4 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Maintaining an effective feedback channel allows users to provide feedback when issues with the bench- mark are discovered
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c8c03501-839a-4f87-beea-c2fbe5820cc7 · outbound
Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models • Justification: Clarifying who holds long- term responsibility reassures the commu- nity that it will be actively supported and improved, ensuring usability and credi- bility
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1117514-0ac2-434e-b5a5-fdffb5bf688e · inbound
Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2d9fc485-8d1a-47c7-a81e-953ad57fe0a4 · inbound
Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extraction Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a063ba73-4ac4-4b98-a403-8fb3e350d976 · inbound
Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b64375e8-78c7-4159-b8ff-7db14240cf13 · inbound
Measuring Epistemic Resilience of LLMs Under Misleading Medical Context Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f54b1aa8-b540-408b-abe8-e49fdb511d9f · inbound
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2b6f817b-7124-4c0f-948e-1322b3e4f9e7 · inbound
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4772a3c0-f726-4610-868e-a35c4e2e6e0f · inbound
MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Models Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0b7158f-dffb-4080-8b8c-4a42b6591ebf · inbound
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cce5bdbf-ff82-4873-b6f3-3f0aba30d6eb · inbound
LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.