Pith. sign in

Paper Citation Record · LEDGER

Collaboration among Multiple Large Language Models for Medical Question Answering

As of 8 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2505.16648.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16648 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:08.628751Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c796a62-7f5a-42ac-8c29-6e6b8be5dbec · outbound

This paper cites Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches.

Collaboration among Multiple Large Language Models for Medical Question Answering Med42 -- Evaluating Fine-Tuning Strategies for Medical LLMs: Full-Parameter vs. Parameter-Efficient Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.621356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.621356Z digest=sha256:48c0dd3da4bbdb650959a6db0aa1140e060957735f9c0af35092f1163198f1e7

Observation 992ffb51-d662-4a81-ab2b-6ca34058509c · outbound

This paper cites Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding.

Collaboration among Multiple Large Language Models for Medical Question Answering Clinical Camel: An Open Expert-Level Medical Language Model with Dialogue-Based Knowledge Encoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.672437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.672437Z digest=sha256:4cf42cec288abbd75f9502995334ad6ffe8629cf33f7e514accf4c0f6715a4fd

Observation 0f4fa892-563d-481b-9ffc-1e0554f68c77 · outbound

This paper cites Meditron-70b: Scaling medical pretraining for large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Meditron-70b: Scaling medical pretraining for large language models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.394863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:05.831701Z digest=sha256:6b98128337a3fe445b6f5dae104df89e78a8ecc60d3f862d723780bcc7dc703a

Observation 0d17389c-b692-49f5-88cb-7b9f5c11f9de · outbound

This paper cites MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data.

Collaboration among Multiple Large Language Models for Medical Question Answering MedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:05.922781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:05.922781Z digest=sha256:6ef897599b60b187988c87536b22e939e1ea3e9a75f8e45baa4b6449080f19ab

Observation 34e413f6-317e-43a8-ba03-a49903e0d401 · outbound

This paper cites ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?.

Collaboration among Multiple Large Language Models for Medical Question Answering ChatGPT and the clinical informatics board examination: the end of unproctored maintenance of certification?

Reference 5

Resolution
verified exact
doi, observed 2026-08-07T15:02:09.035183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:05.989565Z digest=sha256:55bcdaadd5d78268ce84c144044263647815b3e7acbae863ff4f3ff6aeb5f211

Observation 6a53e10c-b5f0-41f5-bd94-620c917bf9bd · outbound

This paper cites Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,.

Collaboration among Multiple Large Language Models for Medical Question Answering Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.105502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.105502Z digest=sha256:3bd086ca2c5fc867d8b2c0a1a2c65753eac05e277291b13c174898492d53e982

Observation 7ca43031-3112-4a53-8b85-2c4bfbcbf81e · outbound

This paper cites How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,.

Collaboration among Multiple Large Language Models for Medical Question Answering How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:13.138444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.175193Z digest=sha256:27d3f1c74b2878e69640d75d675879e9e8baa28f94078c45a9cbe44c4ace85ef

Observation ccf5b023-eedb-4ae1-9c7c-34151049bf1e · outbound

This paper cites Capabilities of GPT-4 on Medical Challenge Problems.

Collaboration among Multiple Large Language Models for Medical Question Answering Capabilities of GPT-4 on Medical Challenge Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:06.263833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:06.263833Z digest=sha256:7f71d37ca5a073674d68d97f407268460c71cb8d73b18bb82778ce824652c2ef

Observation 7ea434a5-7605-427c-a156-480504eb66f0 · outbound

This paper cites Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,.

Collaboration among Multiple Large Language Models for Medical Question Answering Trialling a Large Language Model (ChatGPT) in General Practice With the Applied Knowledge Test: Observational Study Demonstrating Opportunities and Limitations in Primary Care,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.876540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.367380Z digest=sha256:76cfa1bdd11ef2d53ed6cdb2bbeec8eeb546ab9c135146195e1aecfa6f82366c

Observation 12c0d2f2-0c9d-4323-b9dc-397be47b9bcf · outbound

This paper cites Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?.

Collaboration among Multiple Large Language Models for Medical Question Answering Can Large Language Models Safely Address Patient Questions Following Cataract Surgery?

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.367099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.565073Z digest=sha256:86025ed6dd2623ff48260f1445b724c883919178687ca7e72dc176145054b7b4

Observation 515bf289-3aec-4d6a-a9d2-ddba1ebcb442 · outbound

This paper cites Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,.

Collaboration among Multiple Large Language Models for Medical Question Answering Benchmarking large language models’ performances for myopia care: a comparative analysis of ChatGPT-3.5, ChatGPT-4.0, and Google Bard,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.095048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.674469Z digest=sha256:040fcd716f88d493f147c928d396cfd5ed2481033da5ed0087b980b0fbda615a

Observation 611d8b7f-2510-46db-a5ed-78cab0fc3ae5 · outbound

This paper cites Reasoning with large language models for medical question answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering Reasoning with large language models for medical question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.424555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.797610Z digest=sha256:9594f30e1f1bd12b7dbe76595a3f4a4e16ad6105bd8a26a18c6b6f4c9a33cc89

Observation fff2947a-95ac-4d12-944c-032250e05372 · outbound

This paper cites LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,.

Collaboration among Multiple Large Language Models for Medical Question Answering LARGE LANGUAGE MODELS CANNOT SELF- CORRECT REASONING YET,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:11.140056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.843950Z digest=sha256:effdc47b7e20152bae84f4fa920003e1d96d39ab7cac94138d6b42c15fe87172

Observation ac4e2b73-7c80-4877-9b78-f98ca099b073 · outbound

This paper cites Adaptive Chameleon or Stubborn Sloth:,.

Collaboration among Multiple Large Language Models for Medical Question Answering Adaptive Chameleon or Stubborn Sloth:,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.763553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.901674Z digest=sha256:42b7d42b819995952716701cf01804fbfeb5169f3f518d641c865c4eb0e66d21

Observation e06c1b25-ab5e-43bc-b541-fdd8c8d5340c · outbound

This paper cites Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don’t Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.423552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:07.022628Z digest=sha256:06db70902baee0714df5fb51d1f1ed3b278f47495c873ec83a5111c4bd213a9d

Observation fa868ea9-770b-4a39-a717-4ac986eca2ab · outbound

This paper cites Exploring collaboration mechanisms for LLM agents: A social psychology view,.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring collaboration mechanisms for LLM agents: A social psychology view,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:10.087397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:07.118020Z digest=sha256:c0c245aea12018cca69c49b435c32bf4262aea2a6ea570aa11e0ef21f7e267e5

Observation 68489507-5e2d-4677-83fc-7b14593f6f49 · outbound

This paper cites Counterfactual debating with preset stances for hallucination elimination of llms,.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual debating with preset stances for hallucination elimination of llms,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.739336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:07.219238Z digest=sha256:58dfb78e17975c72b41b470330be6da28d74f95ca397e36cfdbce649d5a82e8f

Observation ccb093c0-cc84-469f-bdda-bbdd8eea0de8 · outbound

This paper cites One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,.

Collaboration among Multiple Large Language Models for Medical Question Answering One LLM is not Enough: Harnessing the Power of Ensemble Learning for Medical Question Answering,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.388081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.388081Z digest=sha256:41d96d9c69f0f55513dc00d672873befdc35865eeacd7cba423187cdb609afcc

Observation 20c29ee5-b7aa-46b5-9836-c08769d9753d · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.471313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.471313Z digest=sha256:d95730553b3244db1e064614fc4c1fd5ef9eab91ca23429fca30d9022cf1641c

Observation 9191cfbb-6da6-4588-a62d-48da69099284 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Collaboration among Multiple Large Language Models for Medical Question Answering Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.588376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.588376Z digest=sha256:e3eeb5301c2ebd64bbf1f29f92c13b454e4a2c83cfd06968a364b71eb5cb488a

Observation ecbd8b30-c832-4d9d-ad3e-668d19d1b922 · outbound

This paper cites Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,.

Collaboration among Multiple Large Language Models for Medical Question Answering Multivariable analysis of factors associated with USMLE scores across U.S. medical schools,

Reference 21

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.921281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:07.670413Z digest=sha256:9dc765f31bb39dd6e5760ed4ececa964cdffdd89a617435ce5d79e81541d3963

Observation b6e9e6ad-a41b-409a-8d67-e174dd393f44 · outbound

This paper cites Mixtral of Experts.

Collaboration among Multiple Large Language Models for Medical Question Answering Mixtral of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.767358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.767358Z digest=sha256:b5593b5fa949cd532c948196a4419eaa89478495bfdaddb26b18958769fef0ce

Observation b7d0c431-08c2-4769-8804-95373363d289 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.897800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.897800Z digest=sha256:e98bc6ccd63e4ff666820334c92d72b4acdcca7c4699b185e3913dd6cfbdb80a

Observation 047e65b5-0d3b-449b-b67d-3212a72c01e7 · outbound

This paper cites QLoRA: Efficient Finetuning of Quantized LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering QLoRA: Efficient Finetuning of Quantized LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.959967Z digest=sha256:14d9401ad1367c1122bc084303376b05578feb539b074f62cd8cd4ea1f1638ab

Observation d4d06f39-afbc-4e11-947d-80b83b9585f3 · outbound

This paper cites Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting.

Collaboration among Multiple Large Language Models for Medical Question Answering Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.057788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.057788Z digest=sha256:4222a9d8d5bdf1712632eb0e952e57813f4c839630e4f42506b5daab89a7a7ba

Observation 81563bc2-fe86-4676-b6f6-0f29cb59e984 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Collaboration among Multiple Large Language Models for Medical Question Answering Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:09.405653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:08.188443Z digest=sha256:c713cc53c57f10023fa775ad383ca804a5e2503308feaab221cc1be3efed84c0

Observation 9afeee92-4d2a-42f6-8a53-c48a442c7e24 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

Collaboration among Multiple Large Language Models for Medical Question Answering Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.244822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.244822Z digest=sha256:adaa35856596ca1e49880da8cfafe40d2582a1fe0524767035c33025e4e62de7

Observation 0cdb15c5-516d-4ae5-9db1-7d8b4c6375d9 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Collaboration among Multiple Large Language Models for Medical Question Answering Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.351081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.351081Z digest=sha256:4366d4ca298f87954a7d1bb801892a49c50561a6379f49129752a5e1c1458ce1

Observation f9b95562-1d14-488c-9cf1-2d86e06e9841 · outbound

This paper cites Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View.

Collaboration among Multiple Large Language Models for Medical Question Answering Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology View

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.418344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.418344Z digest=sha256:1d593ffbdcb1e6c6d9f55f36c8a9498efdf4acef01a8a3972834451e3dd439a3

Observation 079fbce1-db73-41f1-a41f-886d28c166c2 · outbound

This paper cites an unresolved cited work.

Collaboration among Multiple Large Language Models for Medical Question Answering Unresolved cited work

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T15:02:08.821707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:08.513674Z digest=sha256:e57436d63cdaa2bec15333055e171fa67d3ec01082a48fd82e6e0597fb4744b1

Observation e20efab2-1a1e-4050-935d-c6177babd4b7 · outbound

This paper cites Debating with More Persuasive LLMs Leads to More Truthful Answers.

Collaboration among Multiple Large Language Models for Medical Question Answering Debating with More Persuasive LLMs Leads to More Truthful Answers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:08.628751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:08.628751Z digest=sha256:6159a446d25c9993ec6115536760ec1b95cdfde4dbbb96429270fd61da1773cc

Observation bae281e0-84d8-4ae1-ae98-baab67d6add2 · outbound

This paper cites Available: https://mededu.jmir.org/2023/1/e46599.

Collaboration among Multiple Large Language Models for Medical Question Answering Available: https://mededu.jmir.org/2023/1/e46599

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:02:12.653097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:02:06.462338Z digest=sha256:82941b308179d6106b58c7d0970dabbd0feaea49cb877f29cb72caa9dcdcf93c

Observation b297728b-67e7-4b33-8e62-14aaf4bb33f1 · outbound

This paper cites Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs.

Collaboration among Multiple Large Language Models for Medical Question Answering Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:07.316036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:07.316036Z digest=sha256:c4164fc4fa105a7020e17f4ea7e124e92f143896508beb052a505243e53cbe19

Pith citing papers

No inbound Pith citation observations are available.