Pith. sign in

Paper Citation Record · LEDGER

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

As of 17 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 6 inbound Pith citation observations for arXiv:2505.11613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11613 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:56:15.821092Z

measured 92 of 92 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:57:11.794516Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T04:19:34.983493Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact3
  • verified fuzzy33
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6804edb4-3393-4e60-a42a-2f200370d63b · outbound

This paper cites write newline.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.471693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.471693Z digest=sha256:f482930dea4cbd759a66eeec9da2113807e1d2dd8e5e43e3e713359f2a5b6d80

Observation 16dd1320-036d-4a6a-bce1-b6cb49a777a7 · outbound

This paper cites GPT-4 Technical Report.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.477495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.477495Z digest=sha256:e6b2f1ee972699ab4e0585bacb90d56bd823f765fe50c212494e492b2f90179a

Observation e0704c29-7930-4e5a-92bd-6cf94d067792 · outbound

This paper cites Publicly Available Clinical BERT Embeddings.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Publicly Available Clinical BERT Embeddings

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.481846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.481846Z digest=sha256:05401ea3c91857415ddc7a4d8ded4e9387590b7cfcd3ecc8e103dcf61689ca2f

Observation a8a9e4c4-f2e1-4f40-91b2-977794fca095 · outbound

This paper cites Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 10 2024.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku, 10 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.486171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.486171Z digest=sha256:7111008cde04db795ab00314d444891762897e5f2d243c71f4c6829281dede77

Observation e31cdd5a-beb8-4e0d-a01c-fcac3dfece1c · outbound

This paper cites Claude 3.7 sonnet, 2 2025.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Claude 3.7 sonnet, 2 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.490173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.490173Z digest=sha256:1e552c8979ae551e17bfd13e76d98ad5cf9c3a6138175fcf78f1c77b34fd10e4

Observation ab949469-ae12-4e3f-925b-70907c2d2cee · outbound

This paper cites an unresolved cited work.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.494220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.494220Z digest=sha256:0ec5908c2477f1cdcd9df2adcdbcfc0320b673a6b0d1713976dd39bfef7167f8

Observation 954b8358-7a3d-4754-8802-5aaf54b48113 · outbound

This paper cites Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.498222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.498222Z digest=sha256:a140beef81c9f1589adb88820f4e4685719076c62d499c160fc1210f934ddb0d

Observation 9d002b8f-407c-41f0-a94a-a9e83ea57162 · outbound

This paper cites Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Large Language Model-informed ECG Dual Attention Network for Heart Failure Risk Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.502724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.502724Z digest=sha256:223b4b9029540a72e9f9e01bf735093eba351b7eb5af5a567ed3465ec3c8c7a1

Observation 485dcf97-659c-483c-92df-11186d6b9006 · outbound

This paper cites MLLM -as-a- Judge : Assessing Multimodal LLM -as-a- Judge with Vision - Language Benchmark.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MLLM -as-a- Judge : Assessing Multimodal LLM -as-a- Judge with Vision - Language Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.506725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.506725Z digest=sha256:a2c2f0b388cc0fe35782d773692e9bcc2def34b9f306f2a89ff31be095d17b03

Observation 3d50859e-408c-4c95-9d1a-8ad184129688 · outbound

This paper cites MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MEDITRON-70B: Scaling Medical Pretraining for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.510739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.510739Z digest=sha256:08ce6aecacd4d66167ef271e69556d691c337cd0fd3fcfe8df1559b5bde9e23e

Observation 1207e76b-521f-4db2-89bf-b817ba4199f9 · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.514883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.514883Z digest=sha256:ee4a0397866459955c9f719c77b3887f8c72d4b44cb622c57757f43d3fc39b1e

Observation 3b78eb87-1c95-463d-b434-0909644f106b · outbound

This paper cites Can LLM Be A Personalized Judge ? arXiv preprint, 2024.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Can LLM Be A Personalized Judge ? arXiv preprint, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.518758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.518758Z digest=sha256:76851fde065f947649088b4f5f6800eeceb86f5dba4b2a7dabc27cc849d4ff37

Observation fa8a0989-c189-46a5-8566-46487a7408bd · outbound

This paper cites Integrating physician diagnostic logic into large language models: Preference learning from process feedback.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Integrating physician diagnostic logic into large language models: Preference learning from process feedback

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:17.051921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.522508Z digest=sha256:d498df803075b7b3e12247dae29f2b2fafdf96b6a861ccdcceec7ea931806b96

Observation 5efbc7a9-a885-441c-865e-a6457e063671 · outbound

This paper cites Autonomous medical evaluation for guideline adherence of large language models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Autonomous medical evaluation for guideline adherence of large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:17.039351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.526263Z digest=sha256:478b728783e4a52eaf179ab452e0cee82c740067ed30ae8a2d089121f1d25b2e

Observation 042ab4f0-a78e-4c4e-bc34-24a97f0834e9 · outbound

This paper cites an unresolved cited work.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:56:17.026633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.529771Z digest=sha256:f6ba082d1a74e2420d8911f2ff10e21a39eec1d1b948f5bd182e21d48500451f

Observation 744c3363-c45a-4897-8357-72933a2a7931 · outbound

This paper cites Lai, Mark J Pletcher, and Ki Lai.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Lai, Mark J Pletcher, and Ki Lai

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:17.013544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.533516Z digest=sha256:6e614045bef1cd5efb5e3c603cbeef0c2264307c79be60e01a542b6fcec0478f

Observation 5bc4e823-6a44-437c-a3c2-d7311cc40e26 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Improving alignment of dialogue agents via targeted human judgements

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.537449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.537449Z digest=sha256:d9515903af213dfb955b20e194c2ac0c301ea20254e69d91cd869258ccad5c69

Observation 51302fb5-7952-4b62-aaf6-ccb024ca8953 · outbound

This paper cites Gemini 2.5 flash: Speed and value at scale, 4 2025.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Gemini 2.5 flash: Speed and value at scale, 4 2025

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:17.000899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.541670Z digest=sha256:7dcaaac5b8684c01d67e2aa9576462bb59bf11c643cdd7d37b02533341788bf7

Observation 0881f608-895f-47e0-ad7d-f27df06f8385 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.545325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.545325Z digest=sha256:914427e365fb58d7884dae9e984e3473aa196e6771a1463bbe09d0236d34e202

Observation 7a471c64-c0c8-4aaf-865c-dc8098a2d926 · outbound

This paper cites Evaluation and mitigation of the limitations of large language models in clinical decision-making.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluation and mitigation of the limitations of large language models in clinical decision-making

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.549317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.549317Z digest=sha256:5e470e22489113d995ad4a9814174cd34c76f86526ead9e662c21ad7612edf7e

Observation 56d26067-15ef-417f-a393-502571d6b2e9 · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.553375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.553375Z digest=sha256:f3570afa073667f0417d7f4305993096778214b3a1858e2e7a78842d07e68161

Observation 78213075-504f-4a20-a9fa-7938e99ae92a · outbound

This paper cites Bp4er: Bootstrap prompting for explicit reasoning in medical dialogue generation.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Bp4er: Bootstrap prompting for explicit reasoning in medical dialogue generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.987415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.557579Z digest=sha256:ad5d0cdeb5f4bce75e7a699d99cbb872d057a3374b2585c9b25f5b607631e7ca

Observation 8e5b63c3-2f74-4b5a-be3c-a50b691edc16 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Measuring Massive Multitask Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.561365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.561365Z digest=sha256:83335cb12cc95d5b594a1786b2b8ddeee89350f300977aa22c3fb7dcd7812b6e

Observation f846b5d1-7e45-4124-9b79-0f65ea53e3e6 · outbound

This paper cites A Benchmark for Long-Form Medical Question Answering.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Benchmark for Long-Form Medical Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.565317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.565317Z digest=sha256:f87af93c91e2c32033af34f5969d108dec72397224bfd982b56ae7b95348caec

Observation bd03f8d1-5111-48e9-876b-6cca31e71b40 · outbound

This paper cites An Empirical Study of LLM -as-a- Judge for LLM Evaluation : Fine -tuned Judge Model Is Not A General Substitute for GPT -4.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models An Empirical Study of LLM -as-a- Judge for LLM Evaluation : Fine -tuned Judge Model Is Not A General Substitute for GPT -4

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.972472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.569524Z digest=sha256:64f308d4f326508164c3b332bfdab01b13808742028ad7f5511fca1db46ebd63

Observation 14b9d893-309f-481b-9326-4a9eb6fe8e25 · outbound

This paper cites A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Comprehensive Survey on Evaluating Large Language Model Applications in the Medical Industry

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.573680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.573680Z digest=sha256:0b802a05496eebf25dc65a100cc29a8f73cd8de274e68bf24752c66134925fc7

Observation 242a90e4-d5ab-4033-980b-55b1845c3137 · outbound

This paper cites What is instruction tuning? IBM, 2023.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models What is instruction tuning? IBM, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.959303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.577693Z digest=sha256:8e51fd5cf70292a50bdfa8ec3ae20754e3e08e4d70a92d99100197dabb4bdc90

Observation 87833239-bbd5-4183-9630-fc6a366eda27 · outbound

This paper cites Mistral 7B.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Mistral 7B

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.581558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.581558Z digest=sha256:498f513469b7643c32d35c19820de3ad680b76dd20fe8484b271e0fb66482120

Observation b9573501-6b35-4987-9325-c2badcbb2e01 · outbound

This paper cites Mixtral of Experts.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Mixtral of Experts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.585609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.585609Z digest=sha256:d04f5e06a43604529827951384308fc21c54d6f9e0fa25181c6c80e2f893d444

Observation 89153992-a290-4710-932f-f96bb1125407 · outbound

This paper cites Health system-scale language models are all-purpose prediction engines.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Health system-scale language models are all-purpose prediction engines

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.946992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.589581Z digest=sha256:72676cab080e64f24b98143944d71d606d75ff321480e30582dac32aadaaba06

Observation c2df650e-d4db-4116-9893-17972f6c5b63 · outbound

This paper cites Pubmedqa: A dataset for biomedical research question answering.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Pubmedqa: A dataset for biomedical research question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.934145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.593507Z digest=sha256:769cc988fc979b2d4dd4d65bc6e3d5ad97045bf1c5a50fde3d57a33bb8d639f0

Observation 36abe63f-8680-48fa-938e-e0919d6ac2d3 · outbound

This paper cites an unresolved cited work.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:56:16.921246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.597387Z digest=sha256:51e3fd70191aeeac095aff83bd380323d3b809685bdfa0ec9456d657108c3acd

Observation d0053479-ca32-4714-ba08-9a418a7d5427 · outbound

This paper cites Towards a multilingual benchmark for medical knowledge assessment.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Towards a multilingual benchmark for medical knowledge assessment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.601213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.601213Z digest=sha256:a54dca15ae73156fb304f3c8146c97f906ff98ec6fa9b7baaa83a25d177d982c

Observation b1b8446a-9ab2-4a62-bd0c-4e12e9898558 · outbound

This paper cites Biobert: a pre-trained biomedical language representation model for biomedical text mining.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Biobert: a pre-trained biomedical language representation model for biomedical text mining

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.605252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.605252Z digest=sha256:c2168313894580dc5bf272b46fa7877c27542ae24db069a36e7dc2252bf601f3

Observation d8acca3f-f0f9-47f2-bfc6-c2b40b426d1b · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.609310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.609310Z digest=sha256:fcdfa58dc75da754fc4b9f526551737fca7330740d6e881195806c073256a4a7

Observation 1484bfb1-2ed5-43f3-9229-596beab0b8ec · outbound

This paper cites Rule-based data selection for large language models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Rule-based data selection for large language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.613839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.613839Z digest=sha256:4c5126830cd5df0a700f1b391fcfd2cec65519b09d0569bc832315ed85647ef3

Observation f6945ad0-c67b-4e03-96ab-c0f97fa959e1 · outbound

This paper cites Multi-head reward aggregation guided by entropy.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Multi-head reward aggregation guided by entropy

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.617631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.617631Z digest=sha256:8d4e1874434f4ed06121b3b993f48219cf498d7f0aeceb0f294f05a1cd8e9b81

Observation 5d6ef36b-588b-4687-8898-ad4c1ce30ca2 · outbound

This paper cites Data-adaptive Safety Rules for Training Reward Models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Data-adaptive Safety Rules for Training Reward Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.621439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.621439Z digest=sha256:a1b8705f75426d923c88ca9187d7888278d9b462554f9306dc72610ba242f079

Observation 21dcd517-5b2d-4ef4-8957-7a5e3c1b2063 · outbound

This paper cites DeepSeek-V3 Technical Report.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models DeepSeek-V3 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.626075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.626075Z digest=sha256:20110984331906afe911f1986835f488d932c8ae8188975263b08778f728ac17

Observation b6141c47-7b77-4238-9954-ddf8308caf09 · outbound

This paper cites MedCalc-Bench: Evaluating Large Language Models for Medical Calculations.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MedCalc-Bench: Evaluating Large Language Models for Medical Calculations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.630445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.630445Z digest=sha256:ca77d54a52e55f6e6cf8b54c2437e94fb2fc6d69228df8a091ff9b473b4a023d

Observation 0d9d9d37-b211-440d-98e8-39c149398914 · outbound

This paper cites Introducing llama 3.1: Our most capable models to date, 7 2024 a.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing llama 3.1: Our most capable models to date, 7 2024 a

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.900031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.634608Z digest=sha256:c5e44b2a34e203795229c00d3824712ee88a0e7d9bbd52c2fea12a3a0fd25162

Observation 9664a628-cc22-4db8-972c-45cc79afb01b · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 9 2024 b.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models, 9 2024 b

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.887409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.638706Z digest=sha256:58f2db593b3c7d3a2a37920daa399e686b204de941d94c1df94657949d3ebfc1

Observation f2385416-be24-48dd-9391-75b15fe04677 · outbound

This paper cites NCCN Clinical Practice Guidelines in Oncology.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models NCCN Clinical Practice Guidelines in Oncology

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.875011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.642392Z digest=sha256:67d18d618a4eadac241ddbd49c58ed72c3d4f275fa16c1094fc18249d89062a1

Observation b9f1b147-3791-43e1-aeac-21605915cfde · outbound

This paper cites Hello gpt-4o, 5 2024 a.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Hello gpt-4o, 5 2024 a

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.862631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.646521Z digest=sha256:9e5554ef8388cc3d7b68222e5ec5006e30f6e2673ace412f77e8119fe8d5611c

Observation 70ef6ca5-7e4e-42a4-81b1-178fd021ac0d · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence, 7 2024 b.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Gpt-4o mini: Advancing cost-efficient intelligence, 7 2024 b

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.849613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.650416Z digest=sha256:2d93dc09e97a2b172e3fc2e6be88025c7f1c7c8f396bd1c5d7bb1b20a32e1cb6

Observation cabb9ae9-465f-4bc1-a123-084b0d87c183 · outbound

This paper cites Openai o1, 9 2024 c.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Openai o1, 9 2024 c

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.836007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.654313Z digest=sha256:0ed552b5741595ed7ee34418a28bd97def164615b9d64ea60c6d1c98af855a23

Observation dae44cfb-7efc-49ec-91e4-239d222940ac · outbound

This paper cites Introducing gpt-4.1 in the api, 4 2025 a.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing gpt-4.1 in the api, 4 2025 a

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.822558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.658270Z digest=sha256:d64ef87738b5aa752cba7cb9231ecc9c4d2a4f26589013571794ec31f7e4435a

Observation 3e683456-b26d-44d3-9b7c-6456f4448c23 · outbound

This paper cites Introducing o3 and o4-mini, 4 2025 b.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Introducing o3 and o4-mini, 4 2025 b

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.809538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.662069Z digest=sha256:df999ad00aa7baaee4b283d2c112f0e4a0c635db9f4361d651af120ec223404a

Observation e2fcb13c-2557-4f02-8cdf-ee2c594e9cca · outbound

This paper cites Who releases ai ethics and governance guidance for large multi-modal models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Who releases ai ethics and governance guidance for large multi-modal models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.795928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.666005Z digest=sha256:e5e9184482c3403c4745d0e17cc31e12ca5112708a69db81291a20a56b175029

Observation 3d69c521-cb1c-422c-a696-19ea729255bf · outbound

This paper cites Medmcqa: A large-scale multi-subject multi-choice dataset for medical question answering.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Medmcqa: A large-scale multi-subject multi-choice dataset for medical question answering

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.782541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.670026Z digest=sha256:b4753a71272a215df83fd0c54ac5b36ae6f13e6c5374e7f7f641b97730161de2

Observation 4c5b5ee8-2b91-4fb2-800b-48cd854bc75d · outbound

This paper cites Note on regression and inheritance in the case of two parents.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Note on regression and inheritance in the case of two parents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.673763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.673763Z digest=sha256:b3be008f98ee08adccf63a909d82a626bb1e8a3a6929a5584f52513c01d2e2d9

Observation f7271b65-b7a6-46f3-893d-904f178004f3 · outbound

This paper cites Efficient Multi -prompt Evaluation of LLMs.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Efficient Multi -prompt Evaluation of LLMs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.759605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.677821Z digest=sha256:08f4137f54c5d465146fb46b41cc9359227fb4c0b5aadf5b2fe3485504ebccbd

Observation a450913b-e757-4e87-a3ad-436aebc1f628 · outbound

This paper cites Qwen3: Think deeper, act faster, 4 2025.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Qwen3: Think deeper, act faster, 4 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.746610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.682590Z digest=sha256:2350d7b0100df1b13a267eaac26452d095316ed369a20245609a04e346ab8352

Observation 9bd8bdbb-7c9c-4278-9635-cef8751517fd · outbound

This paper cites Evaluating Large Language Models at Evaluating Instruction Following.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluating Large Language Models at Evaluating Instruction Following

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.686533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.686533Z digest=sha256:69ccf0b01d66362cb6c658e6f41c713cc2ee3b6659f0ce401057a7b616662f94

Observation 56f2a08e-1c3e-4c90-b5b4-6a7dbd0c5bba · outbound

This paper cites Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Retrieval Augmented Chest X-Ray Report Generation using OpenAI GPT models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.690689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.690689Z digest=sha256:1c59da298d8a7f8fa69827af4ce709aff16777e5ccfd78f79979a0889a55eae4

Observation 6a3a89a4-79cf-41bf-aeb8-27020e3c454f · outbound

This paper cites A context-based chatbot surpasses trained radiologists and generic chatgpt in following the acr appropriateness guidelines.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A context-based chatbot surpasses trained radiologists and generic chatgpt in following the acr appropriateness guidelines

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.732824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.694976Z digest=sha256:cc52c60d67b829a5a93e789c02a470629dafbde7d3ac304c91c84acf82c31411

Observation f81728d7-1478-4aba-ac87-f1a020f44087 · outbound

This paper cites Wisdom of the silicon crowd: Llm ensemble prediction capabilities match human crowd accuracy.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Wisdom of the silicon crowd: Llm ensemble prediction capabilities match human crowd accuracy

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.718841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.698862Z digest=sha256:8fa68b1b7f998d6bc6e5940317a521df43b3eed4399144d7fca6ea1c78811eea

Observation c1e66fff-7c7c-4324-aff4-79122c1b6ec3 · outbound

This paper cites Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Retrieval-augmented large language models for adolescent idiopathic scoliosis patients in shared decision-making

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.705654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.702794Z digest=sha256:1367a099a37cdfecf8674fc92514bf093fa3e810ac79453fe9a4c83b48ae5237

Observation ada3bde4-b71b-4c52-b4ec-61dd2ddfa654 · outbound

This paper cites BioMegatron: Larger Biomedical Domain Language Model.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models BioMegatron: Larger Biomedical Domain Language Model

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.707315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.707315Z digest=sha256:600fd0c4411a25b91f8be3bac130da5dcf0b72d4960dadddcbed386598861bce

Observation 9f3967d4-8bbd-4d8f-a677-aa357ae4b9a3 · outbound

This paper cites Toward expert-level medical question answering with large language models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Toward expert-level medical question answering with large language models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.711449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.711449Z digest=sha256:bd3955f836c963d55e161226d7822627f244f2bc677f8753ca3d0f815293ea55

Observation 87db93ef-0f67-4ed5-a84e-476020431c48 · outbound

This paper cites The proof and measurement of association between two things.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models The proof and measurement of association between two things

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.683460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.715870Z digest=sha256:9c5935751fdca6f3cc09603a0b9050ce2e13ff66229a262bc44719f15b2ee4ae

Observation d1b000d3-8acf-4752-90b9-d0f2088c6a48 · outbound

This paper cites Llamacare: A large medical language model for enhancing healthcare knowledge sharing.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Llamacare: A large medical language model for enhancing healthcare knowledge sharing

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.669862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.720602Z digest=sha256:4e5cdd2081e8163398b9df71ece1c1bae73976efb8b7751ab1e8bf3a2bba13ae

Observation c64bb6a1-8511-43d9-88b6-48f1f86eee6a · outbound

This paper cites Judging The Judges : Evaluating Alignment and Vulnerabilities in LLMs -as- Judges.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Judging The Judges : Evaluating Alignment and Vulnerabilities in LLMs -as- Judges

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.655121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.724695Z digest=sha256:d150e6d78a8b80ca7f9ba0c9799f5d672e8fb7a1594c6d90a5eab2fc005ac6a6

Observation 1bc0ef48-17b8-45a3-b99d-18db0d893362 · outbound

This paper cites Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.728982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.728982Z digest=sha256:0cfb3e99ed4461c076446b14f5ef374159d5283e36f826b9c6878996c707d71b

Observation ef6fc4c2-277f-4ae2-9f4f-a31ada952c28 · outbound

This paper cites Biomedlm: a domain-specific large language model for biomedical text.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Biomedlm: a domain-specific large language model for biomedical text

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.631648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.732856Z digest=sha256:9aa1710a551144b1bb02a2244b70ef1c44b27821073b43ee9eefb1581ac6e351

Observation 83bc2a0f-7989-4b59-bb95-3aac398bcef3 · outbound

This paper cites ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models ClinicalGPT: Large Language Models Finetuned with Diverse Medical Data and Comprehensive Evaluation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.738378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.738378Z digest=sha256:591a0f610b945911e9dc8b766d52964c2fe5c4a39e7e4f4aa3c4e18305d871eb

Observation 32e0f763-8eaf-4e9a-9d1d-fb2715241b9f · outbound

This paper cites Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Beyond Direct Diagnosis: LLM-based Multi-Specialist Agent Consultation for Automatic Diagnosis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.742563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.742563Z digest=sha256:6aa30ae2409097885de4d30638f4da450215eb0fd58546ecc4e4f07f4424ba31

Observation bf3077e8-f43e-42b0-ba74-4400b494c358 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.746778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.746778Z digest=sha256:2da7ae7aaf56f6e6a0db2979881b365115910c376a676591b2e7f55e4e25ef15

Observation 7157acd0-a351-44f4-8acc-5234d64224d1 · outbound

This paper cites A Novel and Accurate BiLSTM Configuration Controller for Modular Soft Robots with Module Number Adaptability.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models A Novel and Accurate BiLSTM Configuration Controller for Modular Soft Robots with Module Number Adaptability

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:56:16.048032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.751002Z digest=sha256:e3f91f1e76727fa66d5522cc221b8db4e17570d3f0f9c061265422ce0dd6329c

Observation c3096ad2-6dcc-4bbe-9c33-ca790ae47390 · outbound

This paper cites MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models MindMap: Knowledge Graph Prompting Sparks Graph of Thoughts in Large Language Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.754945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.754945Z digest=sha256:affaaab8ac59676ce03bd042da9014759b2cd01660e96f8d3db3b63618f64ff4

Observation a7855680-bee7-49ca-bc68-48a2f8b54174 · outbound

This paper cites CodeUltraFeedback : An LLM -as-a- Judge Dataset for Aligning Large Language Models to Coding Preferences.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models CodeUltraFeedback : An LLM -as-a- Judge Dataset for Aligning Large Language Models to Coding Preferences

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.616465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.759369Z digest=sha256:5a26166f2fdf801859115034f82bd46d13ba6b85011f2c4938bdfaafa48da865

Observation ee0bcad6-8264-4df7-9c1c-7e6d85e8f862 · outbound

This paper cites PMC-LLaMA: Towards Building Open-source Language Models for Medicine.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models PMC-LLaMA: Towards Building Open-source Language Models for Medicine

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.763297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.763297Z digest=sha256:8dc8aaa3e6b829fe5cef65109500e889bb853b9eef7b733eba3622136373d579

Observation b4500773-ae0a-419f-9cc8-620425a9e106 · outbound

This paper cites Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Guiding Clinical Reasoning with Large Language Models via Knowledge Seeds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.767234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.767234Z digest=sha256:088c66485409976d7bc087ea5c03a313f8f92d03061d1e388d7eed218da36a76

Observation 518ed48e-ad76-4bcb-a220-250cdfb7e42a · outbound

This paper cites Medkp: Medical dialogue with knowledge enhancement and clinical pathway encoding.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Medkp: Medical dialogue with knowledge enhancement and clinical pathway encoding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.601759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.771546Z digest=sha256:158c3343d79ecd9978424271457af831b2f8ff5ee4fb3258736aab6854051ef6

Observation f432827b-7336-4448-956d-edb2426153e9 · outbound

This paper cites WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models WiseMind: a knowledge-guided multi-agent framework for accurate and empathetic psychiatric diagnosis

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:56:15.990620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.775507Z digest=sha256:b064fdbff745cb8b4a0d627203b6259dcd066e3a15e22de0e6322214369bc971

Observation 155c2e17-e7ee-4296-a313-a8d49dd36377 · outbound

This paper cites Evaluation of large language model performance on the biomedical language understanding and reasoning benchmark.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Evaluation of large language model performance on the biomedical language understanding and reasoning benchmark

Reference 76

Resolution
verified exact
doi, observed 2026-08-15T20:56:15.858313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.779760Z digest=sha256:88ceeecfe0d3f056a1fe3a2b190a250442c13c4c56abd4b1720cb48eb515bced

Observation 04c1bc10-bb3a-453f-8164-1fc83c770dd2 · outbound

This paper cites Qwen2.5 Technical Report.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Qwen2.5 Technical Report

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.783789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.783789Z digest=sha256:e227d998bf13786d4aaf789d4dd2fd9cbde92350778d47c871d757585b657b3a

Observation 8e3f9312-bf2c-4164-84ed-aa3a060f544b · outbound

This paper cites Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Pediatricsgpt: Large language models as chinese medical assistants for pediatric applications

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.586699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.787926Z digest=sha256:c2f0c8b286a5e3786cc07c75f34eb208dea2bbffbbbe001f7d5ddcea3b16b2ae

Observation 9cd6f549-f27e-4e1f-a55f-9b03ec18a253 · outbound

This paper cites GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.791991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.791991Z digest=sha256:0137bd7e29f249f4ac5f2772dd285d4c7219a9cac9dafc1314e5d5224eb33d08

Observation a411b552-dbea-40cf-8392-445e15b0fc7b · outbound

This paper cites InformGen: An AI Copilot for Accurate and Compliant Clinical Research Consent Document Generation.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models InformGen: An AI Copilot for Accurate and Compliant Clinical Research Consent Document Generation

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:56:15.946516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.796320Z digest=sha256:5a0405161b7638a9f1e5e3453591e55cb4dd6d946b3917cf1b82f003ac025ae2

Observation bbf9a165-ebcd-4a83-9a00-69f48a43665e · outbound

This paper cites Huatuogpt, towards taming language model to be a doctor.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Huatuogpt, towards taming language model to be a doctor

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.572270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.800425Z digest=sha256:4e24969d63fd54e233290d239e7c29d7a3a7f7e372b0398b32d030a4ff4781a6

Observation 72538e50-b86d-4baa-9e1f-4a19b4072ef7 · outbound

This paper cites Infobench: Evaluating instruction following ability in large language models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Infobench: Evaluating instruction following ability in large language models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:56:16.559120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T20:56:15.804671Z digest=sha256:67c021eb6032e97acd0ea2af490620a82da09f633241d82f66973774d585e656

Observation d7dc4396-a313-4e63-89b4-47a0db85ed32 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.808565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.808565Z digest=sha256:6e3610e52b62e11687ec1646fe38af2f058e8bf83b50443e27bbff78e2f2cd8e

Observation 42e3ce04-7bdb-4636-887d-94f5dbf4730d · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.812679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.812679Z digest=sha256:1687487756102cd609b9d8d05104e8d77a0394f71e937ad18f782777c89e91ef

Observation e66ad860-7edd-4710-8e1e-e96193c3beb1 · outbound

This paper cites Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Multimodal Large Language Model driven Radiology Report Generation with Clinical Knowledge Enhancement

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.816973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.816973Z digest=sha256:c66ca5a06888bfc480bc5d5f63f6f23aa5f94a1838bf3d039ca9c452dc2c1f03

Observation fc0f981e-ac87-41b4-8831-d48f589129e4 · outbound

This paper cites EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models EMERGE: Enhancing Multimodal Electronic Health Records Predictive Modeling with Retrieval-Augmented Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.821092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.821092Z digest=sha256:a55d3624a8da72d59bb5dab23474b0c7ead08c265622a78f5e79922dd80b4b11

Pith citing papers

Observation 6f9b53ef-fbf0-4de9-a699-ad373d7d76e5 · inbound

Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach cites this paper.

Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T05:37:35.180994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:37:35.180994Z digest=sha256:06fbccdc754089685b2188c48f2880847fb77b04550f389d1a5c655a9237da70

Observation f046df8d-a272-42c6-97cb-7f0d37ff4dc3 · inbound

Evaluating Large Language Models for Evidence-Based Clinical Question Answering cites this paper.

Evaluating Large Language Models for Evidence-Based Clinical Question Answering MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:57:11.794516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:57:11.794516Z digest=sha256:89921ab8b7035adb71e50acb4f3cc17a8302f1a9c89e14bc5e7e19c6d58ed4c8

Observation 6d3ab3ec-af45-4d34-b7ff-1d1ec7b380a8 · inbound

MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion cites this paper.

MuteBench: Modality Unavailability Tolerance Evaluation for Incomplete Multimodal Fusion MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:40.151419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T16:34:06.867042Z digest=sha256:155fff594f3cce731a1789b827babe4dcf13d6d0bd4909ea907cd38f3184e366

Observation be4ea740-4676-4b41-bc49-40ec55aab8d4 · inbound

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment cites this paper.

LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:24:01.689267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T23:18:59.283834Z digest=sha256:2d7b605e16dc0e1ff4a8ed983c3db656e0875635a18d162b4dba56beb97c4ccc

Observation f2fe8203-186c-4fc1-9efa-371d13fc0556 · inbound

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining cites this paper.

Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T04:19:34.985501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T17:02:30.696295Z digest=sha256:fff8c00ce4b26f816df98bb2297b1f30d072fd65783818b3ddec6a74f9654e8a

Observation 6ca38578-7fe8-46a1-b230-56d3ff57ba66 · inbound

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning cites this paper.

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T04:18:55.515412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:18:55.515412Z digest=sha256:499b21f765a4f3c4a255c607a08905f1657770f2ccde42052c94d0476b16f31c