Pith. sign in

Paper Citation Record · LEDGER

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI

As of 19 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2412.12538.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12538 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:40.479370Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact24
  • verified fuzzy15
  • unresolved16
  • parse uncertain0
  • malformed identifier14
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24cb0844-4f0f-4f33-88ac-e376deb41b9b · outbound

This paper cites The theory and practice of clinical decision-making.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The theory and practice of clinical decision-making

Reference 1

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.911904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.234080Z digest=sha256:f241a107268d60387faf1cc3942998e15318f64d9f383f520c11353504d24b1c

Observation 9b243c56-81b2-4dc8-a0e7-73d3986089c1 · outbound

This paper cites Online health: untangling the web.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Online health: untangling the web

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.065881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.237801Z digest=sha256:a7819ad1728b013434aa8e774316811314b8f4a6b6fadb2dd00967d41dbc6757

Observation a7bcd447-bd82-4f8f-9852-7e57e2e830ef · outbound

This paper cites Changes in rates of autopsy-detected diagnostic errors over time: a system- atic review.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Changes in rates of autopsy-detected diagnostic errors over time: a system- atic review

Reference 3

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.899336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.241317Z digest=sha256:22a56ed63a237b7ef849f2c26521a6b86fbce5d13844faaab95501147b3ec0f8

Observation bf88207d-1150-402f-b218-998ba0348f1e · outbound

This paper cites Types and origins of diagnostic errors in primary care set- tings.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Types and origins of diagnostic errors in primary care set- tings

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.904041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.244648Z digest=sha256:bf04f0f1d5e93ba2a90b3b3abccb52799675b34dde53af8a29213369f1f8f510

Observation 89a52009-6a4e-4761-bb45-1d1ebaa1ca42 · outbound

This paper cites Diagnostic errors in hospital- ized adults who died or were transferred to intensive care.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Diagnostic errors in hospital- ized adults who died or were transferred to intensive care

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.056574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.248400Z digest=sha256:230a33a0a830f8f560bcc7583ae9b625c73cc870a2e50a5439d8284847ce7161

Observation 2fa6ba22-6256-42dd-a736-1182fd4d80a3 · outbound

This paper cites National Academies Press; December 29, 2015.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI National Academies Press; December 29, 2015

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:04:40.254839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.254839Z digest=sha256:a07f57652ff70a93a06ee85006164a624f3aa79935f8f987d515ef1fc983b38c

Observation 2714d868-550c-46df-9ed5-1b115dbecc93 · outbound

This paper cites Diagnostic errors in the emergency depart- ment: a systematic review.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Diagnostic errors in the emergency depart- ment: a systematic review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.047134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.258373Z digest=sha256:ebfbf88bc27ea57c2a2b4f91dd5340072fe25c05b6b01839aa2fda5b8097a630

Observation a798061c-efb4-41c5-b1b9-e38f1f531a62 · outbound

This paper cites The ef- fect of Dr Google on doctor-patient encounters in primary care: a quantitative, observational, cross-sectional study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The ef- fect of Dr Google on doctor-patient encounters in primary care: a quantitative, observational, cross-sectional study

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.036821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.262086Z digest=sha256:d241e6c4aa960622f1ae0334bbce3f7f0c37e76152c4afadb2e4c8723d724ee2

Observation 43e1fa64-fd2b-437c-b367-0f09511ced04 · outbound

This paper cites Health Online 2013.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Health Online 2013

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.025943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.265468Z digest=sha256:33108e3e53d42aa4834f72012c84906217caf05e082a83a0d2e389f30f53a6e4

Observation 14663f59-1ba3-459d-a013-649761906c1a · outbound

This paper cites Assessment of Diagnosis and Triage in Validated Case Vignettes Among Nonphysicians Before and After Internet Search.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Assessment of Diagnosis and Triage in Validated Case Vignettes Among Nonphysicians Before and After Internet Search

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.015974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.269282Z digest=sha256:d218c1ec447301e2ba4c7a50b6bb0b95fc63affde1600c3137e74bf5ee2bff37

Observation 0d93c454-2bf9-407e-bdb5-847be6ba886c · outbound

This paper cites A random- ized controlled trial of online symptom search- ing to inform diagnosis.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI A random- ized controlled trial of online symptom search- ing to inform diagnosis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:42.005354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.272553Z digest=sha256:62444ab177e1c58697bd484211a1df8e7cb0d1c7d6e518a4e30e483af97e1940

Observation b9e3638c-97d2-44cd-b082-ae470adc139f · outbound

This paper cites Benchmarking triage capability of symptom checkers against that of medical laypersons: survey study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Benchmarking triage capability of symptom checkers against that of medical laypersons: survey study

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.994475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.276518Z digest=sha256:0ca6aacffed0f84549fa706c73ee968151fb799260be4b9e59241330e8b51690

Observation 710f9f8e-26f1-4c74-8006-ae30a1754088 · outbound

This paper cites Should you search the Internet for information about your acute symptoms? Telemed J E Health.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Should you search the Internet for information about your acute symptoms? Telemed J E Health

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.985128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.279831Z digest=sha256:111d596243154e096b9482b4654326c170fb783e96e00131981e9d5a7186de48

Observation 26812fdd-1606-410f-940d-23c9b11de63f · outbound

This paper cites Internet health infor- mation seeking and the patient-physician re- lationship: A systematic review.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Internet health infor- mation seeking and the patient-physician re- lationship: A systematic review

Reference 14

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.880693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.283119Z digest=sha256:5cb8ba9a9cd0cfe16b7e35f175318d0bbb74e02864403b4fbdf0cff3d55e7b4d

Observation 7ac2a869-aaa3-43c5-aed3-ff28b84368e7 · outbound

This paper cites KFF Health Mis- information Tracking Poll: Artificial In- telligence and Health Information.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI KFF Health Mis- information Tracking Poll: Artificial In- telligence and Health Information

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.976104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.286426Z digest=sha256:ad3481ddb0434131e98555fa69a7cac6588e7adce8f50b6c27ac4a0d20937d79

Observation 0677d6de-eb3c-4943-9776-f1e3b46c6b3e · outbound

This paper cites Appropriateness of Arti- ficial Intelligence Chatbots in Diabetic Foot Ulcer Management.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Appropriateness of Arti- ficial Intelligence Chatbots in Diabetic Foot Ulcer Management

Reference 16

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.868776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.289872Z digest=sha256:b48e4c9caf36bd138dbde470999a0edfc41cea21d633dd7bd3d09ef6e5459e58

Observation 0e96b7a4-f0fd-493d-b75b-a94fe8370b62 · outbound

This paper cites GPT-based chatbot tools are still unreliable in the man- agement of prosthetic joint infections.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI GPT-based chatbot tools are still unreliable in the man- agement of prosthetic joint infections

Reference 17

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.857043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.293313Z digest=sha256:af1ea795e532e57884955bfdd642426b7237da0181ae2b230bd0125b44508f1c

Observation 1adc0dd6-87fe-4126-a0d9-62910ac2e6b9 · outbound

This paper cites Ethical concerns and re- sponsible use of ChatGPT in healthcare.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Ethical concerns and re- sponsible use of ChatGPT in healthcare

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.297321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.297321Z digest=sha256:de515f7872303429a3d43b90756304ed5e5606bf0feba2676ae6f9db57e38b48

Observation a18cc880-c215-460b-82f1-b33e39f35207 · outbound

This paper cites Protocol For Human Evaluation of Artificial Intelligence Chat- bots in Clinical Consultations.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Protocol For Human Evaluation of Artificial Intelligence Chat- bots in Clinical Consultations

Reference 19

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.845454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.300910Z digest=sha256:065b9ff1a28995b7bd401731d9ccedf0567fe1b1047c5c9fbe431102b7110854

Observation 5fe373b5-0648-4903-be73-8bc2c84a6a27 · outbound

This paper cites Balancing Innovation and Professionalism: The Emerging Role of AI-Powered Chatbots in Medical Consultation.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Balancing Innovation and Professionalism: The Emerging Role of AI-Powered Chatbots in Medical Consultation

Reference 20

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.833996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.304827Z digest=sha256:6526b3a11acbb4993a1a618e2e955df026fd298698f31bd8075c249a4b8ae221

Observation bfcae6b5-20e5-4cca-ac0b-207424b6a651 · outbound

This paper cites an unresolved cited work.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:04:41.965709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.308271Z digest=sha256:34972c727009084cc3e4c42e7cd53fdc841e86b0d35bccb61c96f9b87329d506

Observation 4c95b133-7b86-4826-87a4-16a4eecafb8c · outbound

This paper cites Chatbots in Health Care: Connecting Patients to Information: Emerging Health Technologies.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Chatbots in Health Care: Connecting Patients to Information: Emerging Health Technologies

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.956115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.312074Z digest=sha256:b47052a707cfd33fb6259b43971047bcb8aedb3e94927fbdd59723f3f0d7c0eb

Observation 75fb3de2-39d6-4569-aa40-f3a45bc4a82d · outbound

This paper cites Un- derstanding Large Language Models.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Un- derstanding Large Language Models

Reference 23

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.821538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.315437Z digest=sha256:ba08e19531a19db40a80308580e1c369fae27d2d334eab3379659da8a41a4248

Observation d36c5f61-8a73-4dc2-aac2-f6438fa5b52f · outbound

This paper cites The Breakthrough of Large Language Models Re- lease for Medical Applications: 1-Year Timeline and Perspectives.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The Breakthrough of Large Language Models Re- lease for Medical Applications: 1-Year Timeline and Perspectives

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.945759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.318678Z digest=sha256:f16245afca17ad122b298e5c38c26d3877a66ee7b85fd620fa5d0898cce17223

Observation 63de21ca-d0de-4e47-802b-bb9672b184bd · outbound

This paper cites The future landscape of large language models in medicine.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The future landscape of large language models in medicine

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.325033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.325033Z digest=sha256:ce1b9d92ce30b921b6741b0ec7c69af1661a6ab8a63ffbc639e32fb6e2bd21a1

Observation 2c842f2c-9b41-4f4f-bd19-ddc5520b0456 · outbound

This paper cites Large Lan- guage Models for Therapy Recommendations Across 3 Clinical Specialties: Comparative Study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Large Lan- guage Models for Therapy Recommendations Across 3 Clinical Specialties: Comparative Study

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.789666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.328126Z digest=sha256:7814936d7229ca506f02d7761067f6e89c348dfd77d9a579f58e7fc5af5abece

Observation b2199485-1a10-40d2-b5be-8f795fdb9c57 · outbound

This paper cites Large lan- guage models for science and medicine.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Large lan- guage models for science and medicine

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.331206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.331206Z digest=sha256:bb56fb2a59227cc422c47b4b5484769b4e9e660a69fce1cb4e0f2140b68f8329

Observation 973b0ca4-e401-49b1-a226-45118b571cd2 · outbound

This paper cites Evaluation of large language models in breast can- cer clinical scenarios: A comparative analy- sis based on ChatGPT-3.5, ChatGPT-4.0, and Claude2.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Evaluation of large language models in breast can- cer clinical scenarios: A comparative analy- sis based on ChatGPT-3.5, ChatGPT-4.0, and Claude2

Reference 28

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.770626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.334235Z digest=sha256:e491f067ccbeeea2f2a0ca9783226380e133904224a39cfd579f1bbd23bfe8b8

Observation 45618d10-1025-4240-a8fc-3c5c91b533ce · outbound

This paper cites Towards A Deep Learning Question-Answering Specialized Chat- bot for Objective Structured Clinical Examina- tions.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Towards A Deep Learning Question-Answering Specialized Chat- bot for Objective Structured Clinical Examina- tions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.337343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.337343Z digest=sha256:ec045c865b522b1fdfd53c21e4ed40e40e8901093f80f58f99c679f9c38424a2

Observation eee6763d-a9a7-4a96-b747-5f7b5860f6db · outbound

This paper cites Chatbots for Symptom Screening and Patient Education: A Pilot Study on Patient Acceptability in Autoim- mune Inflammatory Diseases.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Chatbots for Symptom Screening and Patient Education: A Pilot Study on Patient Acceptability in Autoim- mune Inflammatory Diseases

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.759394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.340659Z digest=sha256:af6f1a3abf9378ff768524d66c4e78246288c1824d396254d56d0601ccdc8d14

Observation 6009bf80-7b92-460e-befd-e642cef508d7 · outbound

This paper cites Diagnostic reasoning: where we’ve been, where we’re going.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Diagnostic reasoning: where we’ve been, where we’re going

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.613390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.344352Z digest=sha256:b8cce8fd37064ddbf38477e5d737ff41d32c5b5f59dcb7db8a8c673ccccfac7c

Observation f3dd1e13-c9d4-4190-acad-17172b358571 · outbound

This paper cites Clinical decision making.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Clinical decision making

Reference 32

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.747730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.348586Z digest=sha256:3109e4d3bcbaa234111b09e6a8415de7a1daf392fd548b504f914dcecad04987

Observation cd1a6d0a-f91f-4474-92c8-55948f26e1c3 · outbound

This paper cites what is likely to happen.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI what is likely to happen

Reference 33

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.736412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.352545Z digest=sha256:eca0ab30367e8a223bc5fb31d054d5fb122fd35b616a35685c4da92d3b4f3414

Observation 4cdeaf97-e945-4485-ae67-293b005d7eaf · outbound

This paper cites Medicine information needs of patients: the relationships between informa- tion needs, diagnosis and disease.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Medicine information needs of patients: the relationships between informa- tion needs, diagnosis and disease

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.935926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.356509Z digest=sha256:0b85b17d825d8649cea1352b078514a332aae06750ba57ccf8aa717aa6345ece

Observation 055c057a-a950-4b3b-b5c8-5157b28a3483 · outbound

This paper cites Diagnostic rea- soning prompts reveal the potential for large language model interpretability in medicine.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Diagnostic rea- soning prompts reveal the potential for large language model interpretability in medicine

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:04:40.361407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.361407Z digest=sha256:dd84da3f9f2e51918e8d5cacaace062a78aa3a0992da5c685c984a8906eb612d

Observation e591cbea-df03-4989-8349-28da97b2731e · outbound

This paper cites Evaluation of symptom checkers for self diagno- sis and triage: audit study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Evaluation of symptom checkers for self diagno- sis and triage: audit study

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.365026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.365026Z digest=sha256:d72d599b0e1b94bfa817323feffda5f37cdd30c5be010fdf05a452278f76b814

Observation cbb24dad-af2e-4e96-a01a-b098e2982ea6 · outbound

This paper cites Patients don’t present with five choices: an alternative to multiple-choice tests in assessing physicians’ competence.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Patients don’t present with five choices: an alternative to multiple-choice tests in assessing physicians’ competence

Reference 37

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.709901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.368678Z digest=sha256:7d401c41b15dafaafb9167e5d202cc43631794ee8a72354371d2312544cb7776

Observation 427ff340-1fcf-4a46-ab3f-b30469d3df0e · outbound

This paper cites Am- bient artificial intelligence scribes to alle- viate the burden of clinical documentation.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Am- bient artificial intelligence scribes to alle- viate the burden of clinical documentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.372440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.372440Z digest=sha256:6f0fe5ea3c9e6cb1e40d6cf6498884b76b99d0245b5da5c7a1755cfbb370aaf1

Observation fc9f6a54-7638-4b6a-a7e7-5b6b451b365f · outbound

This paper cites Clinical reasoning assessment meth- ods: a scoping review and practical guidance.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Clinical reasoning assessment meth- ods: a scoping review and practical guidance

Reference 39

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.691831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.376138Z digest=sha256:0984675724891828fff5706e8e1973d93445bc7e419628e3206899d9d5d42da1

Observation a4e38aec-a69a-40bd-a72f-405d2fb075e3 · outbound

This paper cites Accuracy of a gen- erative artificial intelligence model in a complex diagnostic challenge.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Accuracy of a gen- erative artificial intelligence model in a complex diagnostic challenge

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.379522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.379522Z digest=sha256:feb975b65993f5aca94ea357339b775e3bf23ffbd165d419ecc41671444c81f1

Observation ac571d2f-525a-4412-b11a-c5b6d97b8727 · outbound

This paper cites How accurate are digital symptom assessment apps for suggesting conditions and urgency advice? A clinical vignettes comparison to GPs.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI How accurate are digital symptom assessment apps for suggesting conditions and urgency advice? A clinical vignettes comparison to GPs

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:04:40.382775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.382775Z digest=sha256:92509caf61dfe829de78f9b5b759f771e0d8dfe2b1a53b4b87c73a17e3a56901

Observation 8607e81f-2790-463a-b5f6-b2ac4f0f8a4e · outbound

This paper cites Chat- GPT influence on medical decision-making, bias, and equity: a randomized study of clin- icians evaluating clinical vignettes.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Chat- GPT influence on medical decision-making, bias, and equity: a randomized study of clin- icians evaluating clinical vignettes

Reference 42

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.673233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.385862Z digest=sha256:1d900444a229afc31b0af9df53635bb9f3bd8e40ec6184963c1651cd197f00be

Observation dc1a7390-d69a-4a3c-84b2-52debc65709a · outbound

This paper cites Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Com- paring physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.389771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.389771Z digest=sha256:7d206aa6ae080b2fb6477d663cbbeac6bb12aa8f10cab156571deee3489dc180

Observation f9bc0552-4a5e-41fa-b6db-24522e539dee · outbound

This paper cites Patient- centered decision making and health care out- comes: an observational study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Patient- centered decision making and health care out- comes: an observational study

Reference 44

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.662396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.393050Z digest=sha256:9b730355653ead89e0996d55a09cc45c61e5f9458d69787c3fd9aae334932425

Observation 42ee19b5-4fd1-46fa-81bb-10b7ac0b4fe9 · outbound

This paper cites Towards Conversational Diagnostic AI.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Towards Conversational Diagnostic AI

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.396592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.396592Z digest=sha256:fc76c5a2d6a822297376643dc5b10628de47b89d8064e6b01c98f19f0c361da8

Observation eac585b0-1750-446c-99a0-17a18ac866ca · outbound

This paper cites Large language models in medicine: the poten- tials and pitfalls: a narrative review.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Large language models in medicine: the poten- tials and pitfalls: a narrative review

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.400841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.400841Z digest=sha256:6188da878acfbeac8c17656fb6d9954f9798b18c55e8156ca70c22e63273e27d

Observation 5f5da1e5-7172-45a3-a460-8c78d823dfb2 · outbound

This paper cites Appropriateness of cardiovas- cular disease prevention recommendations ob- tained from a popular online chat-based ar- tificial intelligence model.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Appropriateness of cardiovas- cular disease prevention recommendations ob- tained from a popular online chat-based ar- tificial intelligence model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.404393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.404393Z digest=sha256:29ecbe7b20fc76889c8ba14226dd3cf3aa3b8326e508a269ada505d6fe9a29f2

Observation e28c47c0-df78-4ce2-adea-e928c9996aae · outbound

This paper cites Young Adults’ Perspectives on the Use of Symptom Checkers for Self-Triage and Self-Diagnosis: Qualitative Study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Young Adults’ Perspectives on the Use of Symptom Checkers for Self-Triage and Self-Diagnosis: Qualitative Study

Reference 48

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.634031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.408594Z digest=sha256:0adeebf429bb58d1da37267bd68b0cd8955f9d890a8db4451ec349f37c06018e

Observation 1fd72fd9-a891-4ffc-8959-0e2569376031 · outbound

This paper cites Ensuring Fairness in Machine Learning to Advance Health Equity.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Ensuring Fairness in Machine Learning to Advance Health Equity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.412457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.412457Z digest=sha256:64401bcd1e4d7de24ee70776123d25dfa0eecc78735581cb473d27b04054cdb0

Observation 71ffa10d-940d-47a3-bcfc-5518f65702ca · outbound

This paper cites Search Engines vs.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Search Engines vs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.419512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.419512Z digest=sha256:7cc9ce5cfb031ae572affb567079dcb5d7eddf0ff85bf966ebd8f2087bf538ed

Observation 3e65d136-2643-455b-b15d-dfe849169df7 · outbound

This paper cites ‘Next please!’ Psychological and practical consequences of an inconclusive diag- Deep Bhatt, Surya Ayyagari and Anuruddh Mishra December 18, 2024 18 nosis.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI ‘Next please!’ Psychological and practical consequences of an inconclusive diag- Deep Bhatt, Surya Ayyagari and Anuruddh Mishra December 18, 2024 18 nosis

Reference 52

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.606289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.422649Z digest=sha256:c87ff1b373ad59d0bd29eeaf56251edccd90600c633dbbce4c72125a9f555594

Observation 1f4e2446-a7b0-4ea2-ac82-21b2e2f987fc · outbound

This paper cites The use of vignettes for conducting healthcare research.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The use of vignettes for conducting healthcare research

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.925678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.425713Z digest=sha256:58179a9a3060fba12fde082066b2f50f38c32f43b0c1a50b3670b107940950f6

Observation 8fc3ad12-8057-4c48-9584-d398be62be0f · outbound

This paper cites Do case vignettes accurately reflect antibiotic prescription? Infection Control and Hospital Epidemiology.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Do case vignettes accurately reflect antibiotic prescription? Infection Control and Hospital Epidemiology

Reference 54

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.596104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.429368Z digest=sha256:a4339b0117546d44df99249199e9a6c025240b7da4173d6af2213f2321c8eb86

Observation d1b77bf0-81f2-45be-ad9c-7b57ccee2e02 · outbound

This paper cites Vignettes as research tools in global health communication: a systematic review of the literature from 2000 to 2020.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Vignettes as research tools in global health communication: a systematic review of the literature from 2000 to 2020

Reference 55

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.191570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.432765Z digest=sha256:a5cdf908293599b295fe829db92ad37b16bbd22ef5e978d0e1ee250d8ac6fd29

Observation 61ca9b53-730a-4e22-9894-649bab965f17 · outbound

This paper cites Com- munication with Diverse Patients: Addressing Culture and Language.Pediatric Clinics of North America.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Com- munication with Diverse Patients: Addressing Culture and Language.Pediatric Clinics of North America

Reference 56

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.585513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.435818Z digest=sha256:65a5fab9f9e6d6766623ccdea84d46de71a5b6f5954abd21c8d7f43eaad667b6

Observation f10f704b-4dcb-4bd1-81c8-b80571cdb8be · outbound

This paper cites Ex- ploring the role of communication barriers in healthcare.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Ex- ploring the role of communication barriers in healthcare

Reference 57

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.117432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.438933Z digest=sha256:d3655c4566f8ced654bbc8c73b501fff159f347a9efa210246555d2b0624c33d

Observation 13fc7ccf-d03d-46ff-8dfe-d5bdce8742a1 · outbound

This paper cites Recognition of patients with medically un- explained physical symptoms by family physi- cians: results of a focus group study.BMC Family Practice.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Recognition of patients with medically un- explained physical symptoms by family physi- cians: results of a focus group study.BMC Family Practice

Reference 58

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.574462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.442336Z digest=sha256:d4ebe326f242fae4d14a16a1afa268a5ae99c249a9a5c80c68a9ee0f28211c08

Observation afa13638-9e6f-4261-a4a2-e3a7b152cbd3 · outbound

This paper cites Do Patients Under- stand? The Permanente Journal.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Do Patients Under- stand? The Permanente Journal

Reference 59

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.563979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.446313Z digest=sha256:b081f3fd3f7492812f60efe1cc33ef1cd71129cea0227b89312c81ba93c801e8

Observation 77e81c7a-4a07-4266-a6a0-d7c1f371348d · outbound

This paper cites The importance of the history and physical in diagnosis.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI The importance of the history and physical in diagnosis

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:41.915087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.449640Z digest=sha256:4129baf247b046654b33b56f68c4692090289e09860d3c6b15319e4e9b9403a0

Observation cc12c55f-c654-457b-804b-74ee80c7d0d1 · outbound

This paper cites In- ternational variations in primary care physi- cian consultation time: a systematic review of 67 countries.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI In- ternational variations in primary care physi- cian consultation time: a systematic review of 67 countries

Reference 61

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.553263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.456829Z digest=sha256:fe99ac9707beae4507d55f6afbe7f60fbc54d4ece6187d746d02d425a6888976

Observation 27b56a44-5d0b-4b90-b345-0f58e8bdf6e2 · outbound

This paper cites Preva- lence of Occupational Burnout among Resident Doctors Working in Public Sector Hospitals in Mumbai.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Preva- lence of Occupational Burnout among Resident Doctors Working in Public Sector Hospitals in Mumbai

Reference 62

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.542624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.460408Z digest=sha256:e56bdac3bf464d29dbeca05cc9e9947eeea71759170bb7700bb37deaa38c7345

Observation 6bc82e7e-c92e-472d-9ea2-31cb2cd2aded · outbound

This paper cites Eval- uating the Diagnostic Performance of Symp- tom Checkers: Clinical Vignette Study.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Eval- uating the Diagnostic Performance of Symp- tom Checkers: Clinical Vignette Study

Reference 63

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.615526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.464042Z digest=sha256:1021ab35abfb7a7dd07f4e0a24bc620276b2a1d8c9f45ed9a78f9a5ca41db138

Observation 7e00d90b-862a-43e3-afe7-98cdfc837dd7 · outbound

This paper cites A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.467589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.467589Z digest=sha256:81377e5a7cd9d452cafa9b93e8b7074b7b7f22e1d970f413d0b6dcf2ed2160db

Observation 01cb11b8-c896-47ec-bfe2-ebd55352aa07 · outbound

This paper cites Evalu- ation framework to guide implementation of AI systems into healthcare settings.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Evalu- ation framework to guide implementation of AI systems into healthcare settings

Reference 65

Resolution
malformed identifier
doi_truncated, observed 2026-08-11T14:04:40.525790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.471160Z digest=sha256:9803e9cfeda907fdeedcfc5352c50f1c0acd317b8cc99bd6d6004a6bd1219717

Observation 10395d0f-4615-4990-8440-e90080a8797c · outbound

This paper cites What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Med- ical Exams.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI What Disease Does This Patient Have? A Large-Scale Open Domain Question Answering Dataset from Med- ical Exams

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:40.475563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.475563Z digest=sha256:2eda5f799eda404893822c60fab1cbe8c3f1b55e8f8e4090cb65da284ad7f816

Observation c97ad694-867e-4b4f-ae24-c05849296bf7 · outbound

This paper cites WHO and ITU establish benchmarking process for artificial intelligence in health.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI WHO and ITU establish benchmarking process for artificial intelligence in health

Reference 67

Resolution
malformed identifier
no resolver link, observed 2026-08-11T14:04:40.479370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:40.479370Z digest=sha256:83eb62b89b2cb4a22d33a9cf87f7e1639cd7a54aa12c321faa19379f45eaf9ff

Observation 16e4ab7d-9c36-4f4c-8d5d-f2f9a29d2148 · outbound

This paper cites an unresolved cited work.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Unresolved cited work

Reference 173

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.828559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.251591Z digest=sha256:df06656e7405cd3ba9dcd56182f7f7ae46547486a8bfac59cd0976468ec6ad1d

Observation 7a725f8e-d38d-46c4-ab46-fdd111b6205d · outbound

This paper cites an unresolved cited work.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Unresolved cited work

Reference 2014

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:41.043691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.453296Z digest=sha256:8304f34b4007167cb096588aadcf8ff6e53d5688a2879c1fa2eb3566ed9c6d3a

Observation e9b99c61-7c13-4c20-bc36-be9bf952e1a6 · outbound

This paper cites an unresolved cited work.

A Scalable Approach to Benchmarking the In-Conversation Differential Diagnostic Accuracy of a Health AI Unresolved cited work

Reference 2024

Resolution
verified exact
doi, observed 2026-08-11T14:04:40.810100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T14:04:40.321795Z digest=sha256:f593633cbd981b21de56638346cd506d52cbba2a0920e0a3b3fefe1722c32884

Pith citing papers

No inbound Pith citation observations are available.