Pith. sign in

Paper Citation Record · LEDGER

Collaboration among Multiple Large Language Models for Medical Question Answering

As of 20 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2505.16648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16648 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:08.628751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c796a62-7f5a-42ac-8c29-6e6b8be5dbec · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Collaboration among Multiple Large Language Models for Medical Question Answering Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.621356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.621356Z digest=sha256:d2880d2f7a1bc5c1c353a5a497fc317250d32def5684d9ed4845b45b5bc42f0b

Observation 992ffb51-d662-4a81-ab2b-6ca34058509c · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Collaboration among Multiple Large Language Models for Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.672437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.672437Z digest=sha256:35bcd2d2e3e27cf72f926594627fe29df86910606b4389fb9390c9333682cb2d

Observation 0f4fa892-563d-481b-9ffc-1e0554f68c77 · outbound

This paper cites Meditron-70b: Scaling medical pretraining for large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Meditron-70b: Scaling medical pretraining for large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.394863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:05.831701Z digest=sha256:406b39ec187986012194d7d0f83b0659a69d5609f8728b2afd821aee8eceb201

Observation 0d17389c-b692-49f5-88cb-7b9f5c11f9de · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

Collaboration among Multiple Large Language Models for Medical Question Answering MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.922781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.922781Z digest=sha256:6c6fe87c156da6367a53795e387592bbe69f194504f560957d135ada67959b8c

Observation 34e413f6-317e-43a8-ba03-a49903e0d401 · outbound

This paper cites ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?.

Collaboration among Multiple Large Language Models for Medical Question Answering ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:02:09.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:05.989565Z digest=sha256:74eb690ab2d8870eab6a3102c74da47632d99884f85d91f0ee4b426bf584f96b

Observation 6a53e10c-b5f0-41f5-bd94-620c917bf9bd · outbound

This paper cites Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.105502Z digest=sha256:590d8890c956a78d128a1c6769dbfb6bf57272cf88a9310c587df2a53eded197

Observation 7ca43031-3112-4a53-8b85-2c4bfbcbf81e · outbound

This paper cites How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,.

Collaboration among Multiple Large Language Models for Medical Question Answering How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.138444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.175193Z digest=sha256:e521ed84b78886868da71da2f1aac6a3b70e531af7aebc6df6427551c072ed3e

Observation ccf5b023-eedb-4ae1-9c7c-34151049bf1e · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

Collaboration among Multiple Large Language Models for Medical Question Answering Capabilities of GPT-4 on Medical Challenge Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.263833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.263833Z digest=sha256:441383f2b06cc1348936dc0c86900635773484fcb8211978eedfd3107e23d9cd

Observation 7ea434a5-7605-427c-a156-480504eb66f0 · outbound

This paper cites Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,.

Collaboration among Multiple Large Language Models for Medical Question Answering Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.876540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.367380Z digest=sha256:42c6722ef2ac9e11dcd718c88939bec797d5548fda075137eedf3d0094f7eec7

Observation 12c0d2f2-0c9d-4323-b9dc-397be47b9bcf · outbound

This paper cites Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?.

Collaboration among Multiple Large Language Models for Medical Question Answering Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.367099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.565073Z digest=sha256:7d1345c04f753366779169021588635ade191ca6e7a5062f8a2d7bc2e16e6099

Observation 515bf289-3aec-4d6a-a9d2-ddba1ebcb442 · outbound

This paper cites Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,.

Collaboration among Multiple Large Language Models for Medical Question Answering Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.095048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.674469Z digest=sha256:eab4a563218f9dda6eda9b35d77e020138f32dffbe669c14cfcac69180f051af

Observation 611d8b7f-2510-46db-a5ed-78cab0fc3ae5 · outbound

This paper cites Reasoning with large language models for medical question answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering Reasoning with large language models for medical question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.424555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.797610Z digest=sha256:3e1907026781a04e9087a9c2dc7f01d90181f9de02a66fd6eaf80e6a59a1d075

Observation fff2947a-95ac-4d12-944c-032250e05372 · outbound

This paper cites LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,.

Collaboration among Multiple Large Language Models for Medical Question Answering LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.140056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.843950Z digest=sha256:07768043fea11bcb25d72dc493cdc8c512d095b370a7ab1c8174e55dcf778aa7

Observation ac4e2b73-7c80-4877-9b78-f98ca099b073 · outbound

This paper cites Adaptive Chameleon or Stubborn Sloth:,.

Collaboration among Multiple Large Language Models for Medical Question Answering Adaptive Chameleon or Stubborn Sloth:,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.763553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.901674Z digest=sha256:a4aee4074c7341c52b2191bd9f199cb6c80c159ad698ecf2b47fef3ad611d2e3

Observation e06c1b25-ab5e-43bc-b541-fdd8c8d5340c · outbound

This paper cites Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.423552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:07.022628Z digest=sha256:a241a077811d0ea1d71b4be6d51a238ab55a6b676ee9758d87d64a941706e496

Observation fa868ea9-770b-4a39-a717-4ac986eca2ab · outbound

This paper cites Exploring collaboration mechanisms for LLM agents: A social psychology view,.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring collaboration mechanisms for LLM agents: A social psychology view,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.087397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:07.118020Z digest=sha256:1d78a3e4f6594b7760dbeb8e392e1d2f80d91be47fdef413bd52bfa56e6c3271

Observation 68489507-5e2d-4677-83fc-7b14593f6f49 · outbound

This paper cites Counterfactual debating with preset stances for hallucination elimination of llms,.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual debating with preset stances for hallucination elimination of llms,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:07.219238Z digest=sha256:55ae332bb0534726c4d2f441a8f6349fa3d5059e9aa2ee5433af4fbe45fb80d6

Observation ccb093c0-cc84-469f-bdda-bbdd8eea0de8 · outbound

This paper cites One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.388081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.388081Z digest=sha256:ef792007881f85b91629707dc08ad9851c2be5892c73506cc736cff6954d9a0c

Observation 20c29ee5-b7aa-46b5-9836-c08769d9753d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.471313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.471313Z digest=sha256:a7c6dcb96f8fd476c2f00024810261041b553a0c72d7bdf9134c0f71f24efd57

Observation 9191cfbb-6da6-4588-a62d-48da69099284 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Collaboration among Multiple Large Language Models for Medical Question Answering Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.588376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.588376Z digest=sha256:eaadf6fbe086f65f136fc8952b26d9b92556118f73c6df7c02021599d3cc6df3

Observation ecbd8b30-c832-4d9d-ad3e-668d19d1b922 · outbound

This paper cites Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,.

Collaboration among Multiple Large Language Models for Medical Question Answering Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,

Reference 21

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.921281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:07.670413Z digest=sha256:975f2174d9a718ecae5274256c4073b6253ca42342256a4ce69213921728024d

Observation b6e9e6ad-a41b-409a-8d67-e174dd393f44 · outbound

This paper cites Mixtral of Experts.

Collaboration among Multiple Large Language Models for Medical Question Answering Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.767358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.767358Z digest=sha256:450baf4aaf14eebf12697037acb16fc55b55b30ffeeefa43a48d94a33018e829

Observation b7d0c431-08c2-4769-8804-95373363d289 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.897800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.897800Z digest=sha256:176b108b53806ab2a4a44146bc05ab3a3d1fe88e81e6832881a98e2455f724bf

Observation 047e65b5-0d3b-449b-b67d-3212a72c01e7 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering QLoRA: Efficient Finetuning of Quantized LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.959967Z digest=sha256:151955d7fd18e8c9e5a77fe5373828ed066def8921926b0a4eacbd240fad93e5

Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.057788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.057788Z digest=sha256:745887f34ea0405f78e8ac9a4bc03973c0708bca6a0b27b04ba6fefd54dc8353

Observation 81563bc2-fe86-4676-b6f6-0f29cb59e984 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Collaboration among Multiple Large Language Models for Medical Question Answering Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:08.188443Z digest=sha256:b9acd3c4fc0622384adc2d085285707ea3b81bc0a9461a2e1a92667a05c1213b

Observation 9afeee92-4d2a-42f6-8a53-c48a442c7e24 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.244822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.244822Z digest=sha256:36356902487bdbba276cefedf9034e8e0fbb5c676f87e0e49db82e02257c45c5

Observation 0cdb15c5-516d-4ae5-9db1-7d8b4c6375d9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.351081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.351081Z digest=sha256:36053147f1667a52f6c3472478a58615ede5e47b9d1110faff6628416f96c7fe

Observation f9b95562-1d14-488c-9cf1-2d86e06e9841 · outbound

This paper cites Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.418344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.418344Z digest=sha256:68777046c52bd937d762cff21d1b837be6a8f16a2a2060b2c9eaefc64f02d771

Observation 079fbce1-db73-41f1-a41f-886d28c166c2 · outbound

This paper cites an unresolved cited work.

Collaboration among Multiple Large Language Models for Medical Question Answering Unresolved cited work

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.821707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:08.513674Z digest=sha256:0a806a939c0c2901358c7832c6a41cd979846e35f7227d7db0daf4dd3f8dab1b

Observation e20efab2-1a1e-4050-935d-c6177babd4b7 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Collaboration among Multiple Large Language Models for Medical Question Answering Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.628751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.628751Z digest=sha256:62f2ec47ded2c1a32225f1b03ffdab8082b17d82d1f8db1c48296a1fb3d1e34a

Observation bae281e0-84d8-4ae1-ae98-baab67d6add2 · outbound

This paper cites Available: https://mededu.jmir.org/2023/1/e46599.

Collaboration among Multiple Large Language Models for Medical Question Answering Available: https://mededu.jmir.org/2023/1/e46599

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.653097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:02:06.462338Z digest=sha256:584729dcf81404051f7b0fb5ac808c74577509e92b4ead158cc2cdc3dd2527ce

Observation b297728b-67e7-4b33-8e62-14aaf4bb33f1 · outbound

This paper cites Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.316036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.316036Z digest=sha256:01bce01877a56c355dc8c6014aee54fead80e0d3d04662c7362ff4b14545927a

Pith citing papers

No inbound Pith citation observations are available.