Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:56:15.821092Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 6 inbound Pith citation observations for arXiv:2505.11613.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:56:15.821092Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:57:11.794516Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T04:19:34.983493Z
86 of 86 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6804edb4-3393-4e60-a42a-2f200370d63b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16dd1320-036d-4a6a-bce1-b6cb49a777a7 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0704c29-7930-4e5a-92bd-6cf94d067792 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Publicly Available Clinical BERT Embeddings
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a9e4c4-f2e1-4f40-91b2-977794fca095 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 10 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31cdd5a-beb8-4e0d-a01c-fcac3dfece1c · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Claude 3.7 sonnet, 2 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab949469-ae12-4e3f-925b-70907c2d2cee · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 954b8358-7a3d-4754-8802-5aaf54b48113 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d002b8f-407c-41f0-a94a-a9e83ea57162 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485dcf97-659c-483c-92df-11186d6b9006 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MLLM -as-a- Judge : Assessing Multimodal LLM -as-a- Judge with Vision - Language Benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d50859e-408c-4c95-9d1a-8ad184129688 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MEDITRON-70B: Scaling Medical Pretraining for Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1207e76b-521f-4db2-89bf-b817ba4199f9 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Harnessing Multiple Large Language Models: A Survey on LLM Ensemble
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b78eb87-1c95-463d-b434-0909644f106b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Can LLM Be A Personalized Judge ? arXiv preprint, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa8a0989-c189-46a5-8566-46487a7408bd · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Integrating physician diagnostic logic into large language models: Preference learning from process feedback
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5efbc7a9-a885-441c-865e-a6457e063671 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Autonomous medical evaluation for guideline adherence of large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 042ab4f0-a78e-4c4e-bc34-24a97f0834e9 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 744c3363-c45a-4897-8357-72933a2a7931 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Lai, Mark J Pletcher, and Ki Lai
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5bc4e823-6a44-437c-a3c2-d7311cc40e26 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Improving alignment of dialogue agents via targeted human judgements
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51302fb5-7952-4b62-aaf6-ccb024ca8953 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Gemini 2.5 flash: Speed and value at scale, 4 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0881f608-895f-47e0-ad7d-f27df06f8385 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a471c64-c0c8-4aaf-865c-dc8098a2d926 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluation and mitigation of the limitations of large language models in clinical decision-making
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d26067-15ef-417f-a393-502571d6b2e9 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78213075-504f-4a20-a9fa-7938e99ae92a · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Bp4er: Bootstrap prompting for explicit reasoning in medical dialogue generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8e5b63c3-2f74-4b5a-be3c-a50b691edc16 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Measuring Massive Multitask Language Understanding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f846b5d1-7e45-4124-9b79-0f65ea53e3e6 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Benchmark for Long-Form Medical Question Answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd03f8d1-5111-48e9-876b-6cca31e71b40 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models An Empirical Study of LLM -as-a- Judge for LLM Evaluation : Fine -tuned Judge Model Is Not A General Substitute for GPT -4
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 14b9d893-309f-481b-9326-4a9eb6fe8e25 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 242a90e4-d5ab-4033-980b-55b1845c3137 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models What is instruction tuning? IBM, 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 87833239-bbd5-4183-9630-fc6a366eda27 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Mistral 7B
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9573501-6b35-4987-9325-c2badcbb2e01 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Mixtral of Experts
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89153992-a290-4710-932f-f96bb1125407 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Health system-scale language models are all-purpose prediction engines
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c2df650e-d4db-4116-9893-17972f6c5b63 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Pubmedqa: A dataset for biomedical research question answering
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 36abe63f-8680-48fa-938e-e0919d6ac2d3 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d0053479-ca32-4714-ba08-9a418a7d5427 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Towards a multilingual benchmark for medical knowledge assessment
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1b8446a-9ab2-4a62-bd0c-4e12e9898558 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Biobert: a pre-trained biomedical language representation model for biomedical text mining
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8acca3f-f0f9-47f2-bfc6-c2b40b426d1b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1484bfb1-2ed5-43f3-9229-596beab0b8ec · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Rule-based data selection for large language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6945ad0-c67b-4e03-96ab-c0f97fa959e1 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Multi-head reward aggregation guided by entropy
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d6ef36b-588b-4687-8898-ad4c1ce30ca2 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Data-adaptive Safety Rules for Training Reward Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dcd517-5b2d-4ef4-8957-7a5e3c1b2063 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models DeepSeek-V3 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6141c47-7b77-4238-9954-ddf8308caf09 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MedCalc-Bench: Evaluating Large Language Models for Medical Calculations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d9d9d37-b211-440d-98e8-39c149398914 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing llama 3.1: Our most capable models to date, 7 2024 a
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9664a628-cc22-4db8-972c-45cc79afb01b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 9 2024 b
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2385416-be24-48dd-9391-75b15fe04677 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models NCCN Clinical Practice Guidelines in Oncology
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b9f1b147-3791-43e1-aeac-21605915cfde · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Hello gpt-4o, 5 2024 a
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 70ef6ca5-7e4e-42a4-81b1-178fd021ac0d · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Gpt-4o mini: Advancing cost-efficient intelligence, 7 2024 b
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cabb9ae9-465f-4bc1-a123-084b0d87c183 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Openai o1, 9 2024 c
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dae44cfb-7efc-49ec-91e4-239d222940ac · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing gpt-4.1 in the api, 4 2025 a
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3e683456-b26d-44d3-9b7c-6456f4448c23 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing o3 and o4-mini, 4 2025 b
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e2fcb13c-2557-4f02-8cdf-ee2c594e9cca · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Who releases ai ethics and governance guidance for large multi-modal models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3d69c521-cb1c-422c-a696-19ea729255bf · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Medmcqa: A large-scale multi-subject multi-choice dataset for medical question answering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4c5b5ee8-2b91-4fb2-800b-48cd854bc75d · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Note on regression and inheritance in the case of two parents
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7271b65-b7a6-46f3-893d-904f178004f3 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Efficient Multi -prompt Evaluation of LLMs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a450913b-e757-4e87-a3ad-436aebc1f628 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Qwen3: Think deeper, act faster, 4 2025
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9bd8bdbb-7c9c-4278-9635-cef8751517fd · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluating Large Language Models at Evaluating Instruction Following
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f2a08e-1c3e-4c90-b5b4-6a7dbd0c5bba · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3a89a4-79cf-41bf-aeb8-27020e3c454f · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A context-based chatbot surpasses trained radiologists and generic chatgpt in following the acr appropriateness guidelines
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f81728d7-1478-4aba-ac87-f1a020f44087 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Wisdom of the silicon crowd: Llm ensemble prediction capabilities match human crowd accuracy
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c1e66fff-7c7c-4324-aff4-79122c1b6ec3 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ada3bde4-b71b-4c52-b4ec-61dd2ddfa654 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models BioMegatron: Larger Biomedical Domain Language Model
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f3967d4-8bbd-4d8f-a677-aa357ae4b9a3 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Toward expert-level medical question answering with large language models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87db93ef-0f67-4ed5-a84e-476020431c48 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models The proof and measurement of association between two things
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d1b000d3-8acf-4752-90b9-d0f2088c6a48 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Llamacare: A large medical language model for enhancing healthcare knowledge sharing
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c64bb6a1-8511-43d9-88b6-48f1f86eee6a · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Judging The Judges : Evaluating Alignment and Vulnerabilities in LLMs -as- Judges
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1bc0ef48-17b8-45a3-b99d-18db0d893362 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6fc4c2-277f-4ae2-9f4f-a31ada952c28 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Biomedlm: a domain-specific large language model for biomedical text
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 83bc2a0f-7989-4b59-bb95-3aac398bcef3 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32e0f763-8eaf-4e9a-9d1d-fb2715241b9f · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf3077e8-f43e-42b0-ba74-4400b494c358 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7157acd0-a351-44f4-8acc-5234d64224d1 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Novel and Accurate BiLSTM Configuration Controller for Modular Soft Robots with Module Number Adaptability
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c3096ad2-6dcc-4bbe-9c33-ca790ae47390 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7855680-bee7-49ca-bc68-48a2f8b54174 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models CodeUltraFeedback : An LLM -as-a- Judge Dataset for Aligning Large Language Models to Coding Preferences
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ee0bcad6-8264-4df7-9c1c-7e6d85e8f862 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models PMC-LLaMA: Towards Building Open-source Language Models for Medicine
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4500773-ae0a-419f-9cc8-620425a9e106 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518ed48e-ad76-4bcb-a220-250cdfb7e42a · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Medkp: Medical dialogue with knowledge enhancement and clinical pathway encoding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f432827b-7336-4448-956d-edb2426153e9 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 155c2e17-e7ee-4296-a313-a8d49dd36377 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluation of large language model performance on the biomedical language understanding and reasoning benchmark
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 04c1bc10-bb3a-453f-8164-1fc83c770dd2 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Qwen2.5 Technical Report
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3f9312-bf2c-4164-84ed-aa3a060f544b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9cd6f549-f27e-4e1f-a55f-9b03ec18a253 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a411b552-dbea-40cf-8392-445e15b0fc7b · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models InformGen: An AI Copilot for Accurate and Compliant Clinical Research Consent Document Generation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bbf9a165-ebcd-4a83-9a00-69f48a43665e · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Huatuogpt, towards taming language model to be a doctor
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 72538e50-b86d-4baa-9e1f-4a19b4072ef7 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Infobench: Evaluating instruction following ability in large language models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d7dc4396-a313-4e63-89b4-47a0db85ed32 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e3ce04-7bdb-4636-887d-94f5dbf4730d · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Instruction-Following Evaluation for Large Language Models
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e66ad860-7edd-4710-8e1e-e96193c3beb1 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc0f981e-ac87-41b4-8831-d48f589129e4 · outbound
MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f9b53ef-fbf0-4de9-a699-ad373d7d76e5 · inbound
Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f046df8d-a272-42c6-97cb-7f0d37ff4dc3 · inbound
Evaluating Large Language Models for Evidence-Based Clinical Question Answering MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d3ab3ec-af45-4d34-b7ff-1d1ec7b380a8 · inbound
MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation be4ea740-4676-4b41-bc49-40ec55aab8d4 · inbound
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f2fe8203-186c-4fc1-9efa-371d13fc0556 · inbound
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6ca38578-7fe8-46a1-b230-56d3ff57ba66 · inbound
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.