Pith. sign in

Paper Citation Record · LEDGER

Collaboration among Multiple Large Language Models for Medical Question Answering

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2505.16648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16648 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:08.628751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c796a62-7f5a-42ac-8c29-6e6b8be5dbec · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Collaboration among Multiple Large Language Models for Medical Question Answering Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.621356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.621356Z digest=sha256:48c0dd3da4bbdb650959a6db0aa1140e060957735f9c0af35092f1163198f1e7

Observation 992ffb51-d662-4a81-ab2b-6ca34058509c · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Collaboration among Multiple Large Language Models for Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.672437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.672437Z digest=sha256:4cf42cec288abbd75f9502995334ad6ffe8629cf33f7e514accf4c0f6715a4fd

Observation 0f4fa892-563d-481b-9ffc-1e0554f68c77 · outbound

This paper cites Meditron-70b: Scaling medical pretraining for large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Meditron-70b: Scaling medical pretraining for large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.394863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:05.831701Z digest=sha256:504c20c61a97cf74f4cd729de97b1d3e6ce2a7642086ccae840c8ba7b832c328

Observation 0d17389c-b692-49f5-88cb-7b9f5c11f9de · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

Collaboration among Multiple Large Language Models for Medical Question Answering MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.922781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.922781Z digest=sha256:6ef897599b60b187988c87536b22e939e1ea3e9a75f8e45baa4b6449080f19ab

Observation 34e413f6-317e-43a8-ba03-a49903e0d401 · outbound

This paper cites ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?.

Collaboration among Multiple Large Language Models for Medical Question Answering ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:02:09.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:05.989565Z digest=sha256:96ea6622f1023af8ced44b157895c3c27b3161db8bfb6e1661b8c111f146ac7b

Observation 6a53e10c-b5f0-41f5-bd94-620c917bf9bd · outbound

This paper cites Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.105502Z digest=sha256:3bd086ca2c5fc867d8b2c0a1a2c65753eac05e277291b13c174898492d53e982

Observation 7ca43031-3112-4a53-8b85-2c4bfbcbf81e · outbound

This paper cites How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,.

Collaboration among Multiple Large Language Models for Medical Question Answering How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.138444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.175193Z digest=sha256:c3005bb8fbdccf04d4526173251fc9ac5341cada18731429654a6f8787553226

Observation ccf5b023-eedb-4ae1-9c7c-34151049bf1e · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

Collaboration among Multiple Large Language Models for Medical Question Answering Capabilities of GPT-4 on Medical Challenge Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.263833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.263833Z digest=sha256:7f71d37ca5a073674d68d97f407268460c71cb8d73b18bb82778ce824652c2ef

Observation 7ea434a5-7605-427c-a156-480504eb66f0 · outbound

This paper cites Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,.

Collaboration among Multiple Large Language Models for Medical Question Answering Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.876540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.367380Z digest=sha256:bdc5cce1e7afb86a70fb956184f58396a655933418ec180472207debf564a722

Observation 12c0d2f2-0c9d-4323-b9dc-397be47b9bcf · outbound

This paper cites Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?.

Collaboration among Multiple Large Language Models for Medical Question Answering Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.367099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.565073Z digest=sha256:60d314c250245bccb4fc8ea9ce71d1c9c17c98df170622a5a53e93f02e888d72

Observation 515bf289-3aec-4d6a-a9d2-ddba1ebcb442 · outbound

This paper cites Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,.

Collaboration among Multiple Large Language Models for Medical Question Answering Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.095048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.674469Z digest=sha256:e0cdbdcb5ba95372697cc31b06317ae84da6a977ddf3540e31713b9c254bcbcb

Observation 611d8b7f-2510-46db-a5ed-78cab0fc3ae5 · outbound

This paper cites Reasoning with large language models for medical question answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering Reasoning with large language models for medical question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.424555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.797610Z digest=sha256:5698d9f0a72544b12a3a544a04a1835ba4858e072c8bf26194ae1cb8074b596e

Observation fff2947a-95ac-4d12-944c-032250e05372 · outbound

This paper cites LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,.

Collaboration among Multiple Large Language Models for Medical Question Answering LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.140056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.843950Z digest=sha256:f5d38f1b08484109994271bc1bc5cbf1453e00292c51eb9614fa0d099ed267e6

Observation ac4e2b73-7c80-4877-9b78-f98ca099b073 · outbound

This paper cites Adaptive Chameleon or Stubborn Sloth:,.

Collaboration among Multiple Large Language Models for Medical Question Answering Adaptive Chameleon or Stubborn Sloth:,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.763553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.901674Z digest=sha256:e4bec9344c841738142bc741f1feb72006d799edde8a02746abf401eba1d3c0a

Observation e06c1b25-ab5e-43bc-b541-fdd8c8d5340c · outbound

This paper cites Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.423552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:07.022628Z digest=sha256:df9271af895b3553614ac1d0eac6bb03179308ec526da932833fa50f5fbbd842

Observation fa868ea9-770b-4a39-a717-4ac986eca2ab · outbound

This paper cites Exploring collaboration mechanisms for LLM agents: A social psychology view,.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring collaboration mechanisms for LLM agents: A social psychology view,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.087397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:07.118020Z digest=sha256:68fff126c348cd4ac9af2b340da93f7179405a2268aea4ba708e75f49612c965

Observation 68489507-5e2d-4677-83fc-7b14593f6f49 · outbound

This paper cites Counterfactual debating with preset stances for hallucination elimination of llms,.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual debating with preset stances for hallucination elimination of llms,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:07.219238Z digest=sha256:2351710e5a09d4431db3e913c7f035c58945c89008e3797c34c2dcb1ffb45925

Observation ccb093c0-cc84-469f-bdda-bbdd8eea0de8 · outbound

This paper cites One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.388081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.388081Z digest=sha256:41d96d9c69f0f55513dc00d672873befdc35865eeacd7cba423187cdb609afcc

Observation 20c29ee5-b7aa-46b5-9836-c08769d9753d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.471313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.471313Z digest=sha256:d95730553b3244db1e064614fc4c1fd5ef9eab91ca23429fca30d9022cf1641c

Observation 9191cfbb-6da6-4588-a62d-48da69099284 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Collaboration among Multiple Large Language Models for Medical Question Answering Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.588376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.588376Z digest=sha256:e3eeb5301c2ebd64bbf1f29f92c13b454e4a2c83cfd06968a364b71eb5cb488a

Observation ecbd8b30-c832-4d9d-ad3e-668d19d1b922 · outbound

This paper cites Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,.

Collaboration among Multiple Large Language Models for Medical Question Answering Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,

Reference 21

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.921281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:07.670413Z digest=sha256:52c79b8dd2c78ead51a6a75a921c0d74cb6b732baa6c23d99950859c3f9a381e

Observation b6e9e6ad-a41b-409a-8d67-e174dd393f44 · outbound

This paper cites Mixtral of Experts.

Collaboration among Multiple Large Language Models for Medical Question Answering Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.767358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.767358Z digest=sha256:ba50a82485fe73eaab4f97cfb3d4ec923e8c2157d6463e41a264eac360faa079

Observation b7d0c431-08c2-4769-8804-95373363d289 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.897800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.897800Z digest=sha256:e98bc6ccd63e4ff666820334c92d72b4acdcca7c4699b185e3913dd6cfbdb80a

Observation 047e65b5-0d3b-449b-b67d-3212a72c01e7 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering QLoRA: Efficient Finetuning of Quantized LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.959967Z digest=sha256:14d9401ad1367c1122bc084303376b05578feb539b074f62cd8cd4ea1f1638ab

Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.057788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.057788Z digest=sha256:4222a9d8d5bdf1712632eb0e952e57813f4c839630e4f42506b5daab89a7a7ba

Observation 81563bc2-fe86-4676-b6f6-0f29cb59e984 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Collaboration among Multiple Large Language Models for Medical Question Answering Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:08.188443Z digest=sha256:658ce6527462942fa9fe40e370a225900aab818e7bc131f4f67124c0bc53acd4

Observation 9afeee92-4d2a-42f6-8a53-c48a442c7e24 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.244822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.244822Z digest=sha256:adaa35856596ca1e49880da8cfafe40d2582a1fe0524767035c33025e4e62de7

Observation 0cdb15c5-516d-4ae5-9db1-7d8b4c6375d9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.351081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.351081Z digest=sha256:4366d4ca298f87954a7d1bb801892a49c50561a6379f49129752a5e1c1458ce1

Observation f9b95562-1d14-488c-9cf1-2d86e06e9841 · outbound

This paper cites Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.418344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.418344Z digest=sha256:1d593ffbdcb1e6c6d9f55f36c8a9498efdf4acef01a8a3972834451e3dd439a3

Observation 079fbce1-db73-41f1-a41f-886d28c166c2 · outbound

This paper cites an unresolved cited work.

Collaboration among Multiple Large Language Models for Medical Question Answering Unresolved cited work

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.821707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:08.513674Z digest=sha256:50f0132aae056da446a01ac08b3acdbd85cd7ae4d93bdbfae0a3bea4a8bffdf6

Observation e20efab2-1a1e-4050-935d-c6177babd4b7 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Collaboration among Multiple Large Language Models for Medical Question Answering Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.628751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.628751Z digest=sha256:6159a446d25c9993ec6115536760ec1b95cdfde4dbbb96429270fd61da1773cc

Observation bae281e0-84d8-4ae1-ae98-baab67d6add2 · outbound

This paper cites Available: https://mededu.jmir.org/2023/1/e46599.

Collaboration among Multiple Large Language Models for Medical Question Answering Available: https://mededu.jmir.org/2023/1/e46599

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.653097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T15:02:06.462338Z digest=sha256:e36ad9894742ee96f0ea8699e72a9342bdffbb7c7abbe8992f6f7ff8c2ae9aa6

Observation b297728b-67e7-4b33-8e62-14aaf4bb33f1 · outbound

This paper cites Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.316036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.316036Z digest=sha256:c4164fc4fa105a7020e17f4ea7e124e92f143896508beb052a505243e53cbe19

Pith citing papers

No inbound Pith citation observations are available.