Pith. sign in

REVIEW 3 major objections 2 minor 53 cited by

Capabilities of Gemini Models in Medicine

T0 review · 3 major / 2 minor · reviewed 2026-05-15 · grok-4.3

Pith's one-line read Med-Gemini models reach 91.1 percent accuracy on USMLE medical questions and surpass GPT-4 on medical benchmarks.

desk verdict Med-Gemini hits 91.1% on MedQA and widens the multimodal gap over GPT-4V, but the gains rest on clean benchmarks without checks for real clinical noise. read the letter →

arxiv 2404.18416 v2 pith:ML5KRKZ2 submitted 2024-04-29 cs.AI cs.CLcs.CVcs.LG

classification cs.AIcs.CLcs.CVcs.LG
keywords Med-GeminimedicalAImultimodalmodelsMedQAUSMLEGeministate-of-the-artperformancebenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents Med-Gemini as a family of multimodal models built from Gemini foundations and specialized for medicine through web search integration and custom encoders for new data types. It reports new state-of-the-art results on 10 of 14 medical benchmarks, with the top model hitting 91.1 percent accuracy on MedQA using uncertainty-guided search and an average 44.5 percent relative gain over GPT-4V on seven multimodal tasks. A sympathetic reader would care because these results point to AI systems that could soon support diagnostic reasoning, long patient record analysis, medical education, and text summarization at or above expert levels. The work also shows strong in-context performance on needle-in-a-haystack retrieval from health records and medical video question answering without task-specific fine-tuning.

What carries the argument

The Med-Gemini family of multimodal models that add medical specialization to Gemini's core strengths via web search access and custom encoders for novel modalities.

What would settle it

A side-by-side comparison of Med-Gemini outputs against board-certified physicians on a held-out set of real de-identified hospital cases, measuring diagnostic accuracy and rate of unsafe recommendations.

Watch

Extended reading notes

Core claim

Med-Gemini models, specialized for medicine with seamless web search and custom encoders, establish new state-of-the-art performance on 10 out of 14 medical benchmarks. The best model achieves 91.1 percent accuracy on the MedQA USMLE benchmark through a novel uncertainty-guided search strategy, surpasses the GPT-4 model family on every benchmark with direct comparison, and improves over GPT-4V by an average relative margin of 44.5 percent across seven multimodal benchmarks including NEJM Image Challenges and the health subset of MMMU. The models further demonstrate long-context strengths by achieving state-of-the-art results on retrieval from long de-identified health records and medical-vid

Load-bearing premise

High accuracy on curated medical benchmarks will translate to reliable and safe performance when the models encounter noisy, incomplete, or out-of-distribution patient data in actual clinical settings.

Editorial extensions

If this is right

  • AI systems could match or exceed human performance on medical text summarization and multimodal image interpretation tasks.
  • Uncertainty-guided search offers a practical way to raise accuracy on medical question answering without additional training.
  • Long-context capabilities enable effective in-context use of full patient histories and video data for research and education.
  • The same specialization pattern could be applied to other high-stakes domains that need up-to-date knowledge and multimodal reasoning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Real-world deployment would still require separate safety testing on noisy clinical data, since benchmark gains do not automatically guarantee clinical reliability.
  • The 44.5 percent multimodal margin suggests similar gains may appear in other visual-heavy medical workflows such as radiology or pathology.
  • Because the models rely on web search, their outputs could be kept current with new guidelines more easily than static fine-tuned systems.
  • Strong needle-in-a-haystack retrieval performance opens the possibility of using the models to surface relevant prior cases from large hospital archives.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The manuscript introduces Med-Gemini, a family of multimodal models specialized for medicine by building on Gemini's core strengths in multimodal and long-context reasoning. It reports new state-of-the-art results on 10 of 14 medical benchmarks, including 91.1% accuracy on MedQA (USMLE) via a novel uncertainty-guided search strategy, consistent outperformance over the GPT-4 family on all directly comparable benchmarks, and a 44.5% average relative improvement over GPT-4V across 7 multimodal benchmarks (NEJM Image Challenges, MMMU health subset, etc.). Additional results highlight long-context retrieval from de-identified health records, medical video QA, and surpassing human experts on medical text summarization, while noting the need for further rigorous evaluation before clinical deployment.

Significance. If the benchmark results prove robust, the work establishes a meaningful advance in medical AI capabilities by showing how general multimodal models can be efficiently specialized for high-stakes domains. The combination of web search integration, custom encoders, and in-context long-context handling without bespoke training is a notable strength. These results provide concrete evidence of progress toward AI support for medical reasoning, education, and research, while the paper's explicit caution about real-world deployment aligns with the safety-critical context.

major comments (3)
  1. Abstract and results sections: The reported SoTA figures (91.1% on MedQA, 44.5% relative margin on multimodal tasks) are presented without error bars, confidence intervals, or multiple-run statistics. This omission makes it impossible to determine whether the gains over prior methods are statistically significant or sensitive to random seeds, directly affecting the reliability of the central performance claims.
  2. MedQA evaluation: The uncertainty-guided search strategy is described as key to reaching 91.1% accuracy, yet no ablation studies compare it against standard search or report its contribution in isolation. Without these controls, it remains unclear whether the performance stems from the strategy itself or from other unstated factors such as prompt engineering or data access.
  3. Multimodal benchmarks section: The average 44.5% relative improvement over GPT-4V is given across 7 tasks, but per-benchmark breakdowns with variance or standard deviations are not provided. This prevents assessment of whether the gains are consistent or driven by a subset of easier tasks, weakening the broad outperformance claim.
minor comments (2)
  1. The abstract could explicitly list the exact prior SoTA scores and the specific GPT-4 variants used for each direct comparison to improve transparency.
  2. Terminology such as 'needle-in-a-haystack retrieval' in the long-context experiments would benefit from a short definition or citation on first use for broader accessibility.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their constructive comments, which highlight important aspects of result presentation and evaluation rigor. We address each major point below and indicate the revisions planned for the next version of the manuscript.

read point-by-point responses
  1. Referee: Abstract and results sections: The reported SoTA figures (91.1% on MedQA, 44.5% relative margin on multimodal tasks) are presented without error bars, confidence intervals, or multiple-run statistics. This omission makes it impossible to determine whether the gains over prior methods are statistically significant or sensitive to random seeds, directly affecting the reliability of the central performance claims.

    Authors: We agree that error bars and statistical measures would strengthen the presentation. The reported numbers reflect single-run evaluations on fixed benchmarks, which is standard practice for large multimodal models given the high computational cost of repeated full evaluations. In the revision we will add explicit discussion of this limitation in the results section, note that gains are consistent with prior single-run reports in the literature, and include any available variance from prompt-sensitivity checks performed during development. revision: partial

  2. Referee: MedQA evaluation: The uncertainty-guided search strategy is described as key to reaching 91.1% accuracy, yet no ablation studies compare it against standard search or report its contribution in isolation. Without these controls, it remains unclear whether the performance stems from the strategy itself or from other unstated factors such as prompt engineering or data access.

    Authors: We will add a targeted ablation in the revised manuscript (main text or supplementary) that isolates the uncertainty-guided search by comparing it directly against a standard search baseline and a chain-of-thought baseline using the identical Med-Gemini model, prompt template, and retrieval setup. This will quantify the incremental contribution of the uncertainty component. revision: yes

  3. Referee: Multimodal benchmarks section: The average 44.5% relative improvement over GPT-4V is given across 7 tasks, but per-benchmark breakdowns with variance or standard deviations are not provided. This prevents assessment of whether the gains are consistent or driven by a subset of easier tasks, weakening the broad outperformance claim.

    Authors: We will expand the multimodal results section to include a full per-benchmark table listing absolute accuracies for both Med-Gemini and GPT-4V on each of the seven tasks, together with the relative improvement for every individual benchmark. Any variance estimates obtainable from the evaluation logs will be reported; otherwise we will note the single-run nature consistently with the first comment. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical benchmark results are self-contained measurements

full rationale

The paper presents Med-Gemini as a family of specialized multimodal models evaluated directly on 14 fixed medical benchmarks (MedQA, NEJM Image Challenges, MMMU health subset, etc.). All reported results—91.1% MedQA accuracy via uncertainty-guided search, 44.5% relative gains over GPT-4V, long-context retrieval, and human-expert comparisons—are obtained by running the models on these curated test sets and recording accuracy or other metrics. No equations, parameter fits, uniqueness theorems, or ansatzes are invoked whose outputs are then relabeled as predictions; the central claims reduce only to the empirical measurements themselves. Self-citations, if present, are limited to prior Gemini work and do not carry the load of the performance numbers. The derivation chain is therefore empty of circular steps.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The work relies on standard large-language-model training assumptions and the premise that public medical benchmarks are representative; no new free parameters, axioms, or invented entities are introduced in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Capabilities of Gemini Models in Medicine." pith.science (2026). https://pith.science/paper/ML5KRKZ2

@misc{pith2026240418416,
  author       = {Pith},
  title        = {Pith review of: Capabilities of Gemini Models in Medicine},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/ML5KRKZ2}},
  note         = {Machine review of arXiv:2404.18416}
}
read the original abstract

Excellence in a wide variety of medical applications poses considerable challenges for AI, requiring advanced reasoning, access to up-to-date medical knowledge and understanding of complex multimodal data. Gemini models, with strong general capabilities in multimodal and long-context reasoning, offer exciting possibilities in medicine. Building on these core strengths of Gemini, we introduce Med-Gemini, a family of highly capable multimodal models that are specialized in medicine with the ability to seamlessly use web search, and that can be efficiently tailored to novel modalities using custom encoders. We evaluate Med-Gemini on 14 medical benchmarks, establishing new state-of-the-art (SoTA) performance on 10 of them, and surpass the GPT-4 model family on every benchmark where a direct comparison is viable, often by a wide margin. On the popular MedQA (USMLE) benchmark, our best-performing Med-Gemini model achieves SoTA performance of 91.1% accuracy, using a novel uncertainty-guided search strategy. On 7 multimodal benchmarks including NEJM Image Challenges and MMMU (health & medicine), Med-Gemini improves over GPT-4V by an average relative margin of 44.5%. We demonstrate the effectiveness of Med-Gemini's long-context capabilities through SoTA performance on a needle-in-a-haystack retrieval task from long de-identified health records and medical video question answering, surpassing prior bespoke methods using only in-context learning. Finally, Med-Gemini's performance suggests real-world utility by surpassing human experts on tasks such as medical text summarization, alongside demonstrations of promising potential for multimodal medical dialogue, medical research and education. Taken together, our results offer compelling evidence for Med-Gemini's potential, although further rigorous evaluation will be crucial before real-world deployment in this safety-critical domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 53 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 93 citations worldwide. Full citation record

  1. Fully Open Meditron: An Auditable Pipeline for Clinical LLMs

    cs.AI 2026-05 unverdicted novelty 8.0 of 10

    Presents the first fully open pipeline for clinical LLMs that unifies eight public QA datasets with clinician-vetted synthetic data from guidelines and vignettes, achieving improved performance on medical benchmarks w...

  2. Aligning Language Models with Selective Prediction

    cs.LG 2026-07 accept novelty 7.0 of 10

    RLSR aligns LLMs via a lifted AURC reward and batch ranking inside GRPO, producing better risk-coverage curves than accuracy- or calibration-based RL on in- and out-of-domain tasks.

  3. MMBU: A Massive Multi-modal Biomedical Understanding Benchmark to Probe the Perception Capabilities of Vision-Language Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    Introduces MMBU benchmark for VLMs in biomedicine and demonstrates that established benchmarks mask perception deficiencies in evaluated models.

  4. DDX-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    DDX-TRACE is a physician-adjudicated benchmark for evaluating VLMs on evidence-supported diagnostic trajectories rather than final answers alone in multimodal neuroradiology.

  5. ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion

    cs.LG 2026-04 unverdicted novelty 7.0 of 10

    ECHO is a one-step block diffusion VLM for chest X-ray reports that improves RaTE and SemScore by over 60% while delivering 8x faster inference than autoregressive baselines.

  6. Revisiting Model Inversion Evaluation: From Misleading Standards to Reliable Privacy Assessment

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Standard model inversion evaluation counts many adversarial false positives as successes; MLLM-based evaluation reveals consistently high false-positive rates across 27 attack setups.

  7. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  8. Evidence-Grounded AI for Musculoskeletal Care

    cs.AI 2026-07 unverdicted novelty 6.0 of 10

    OrthoPilot, an LLM clinical system integrating live hospital data and external knowledge, reportedly beat 25-year orthopaedic experts and raised full-chain management success and bed throughput in multi-site studies.

  9. OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice

    cs.AI 2026-07 conditional novelty 6.0 of 10

    VLMs show a Semantic-Physical Gap on food images: strong dish naming but high MAPE on mass/nutrients and frequent unsafe advice for high-risk disease profiles.

  10. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  11. Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Fine-tuning Gemini 2.5 Pro with LoRA on 400 home videos improves per-feature agreement with clinicians by 40% and zero-shot ASD diagnosis F1 by 53% on held-out data, with classifier pipelines reaching 77% accuracy.

  12. Learning to Recover Task Experts from a Multi-Task Merged Model

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    ReTeX predicts additive offsets to undo merging interference and uses an SVD subspace signature task identifier to recover over 95% of expert performance while improving generalization to unseen tasks.

  13. Cohort-Anchored Foundation Models for Electronic Health Records: From Risk Scores to Auditable Peer Cohorts

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    CAFM is a four-stage framework that anchors EHR foundation models to patient cohorts via deviation-aware curation, cohort-conditioned pretraining, multimodal alignment, and clinician refinement to improve interpretabi...

  14. RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media

    cs.SI 2026-06 unverdicted novelty 6.0 of 10

    RELIANCE is a new expert-annotated dataset of TikTok reproductive health content paired with LLM fact-checking evaluations showing 60% accuracy in sampled videos and a 15% gap between claim and full-content assessment.

  15. Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Med-R2 Bench is a new adversarial benchmark revealing that medical VLMs show sequential performance drops along clinical workflow stages and depend more on correct prompts than on visual grounding.

  16. MoBayes: A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Separating language from Bayesian reasoning lets cheap LLM sensors beat larger standalone LLM doctors on conversational diagnosis with controllable abstention and lower cost.

  17. Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain

    cs.AI 2026-03 conditional novelty 6.0 of 10

    A fixed, sampling-free score — per-token log-probability variance times (1 + average |image-vs-text probability shift|) — detects medical-VQA hallucinations better than semantic-entropy baselines in 13 of 16 settings.

  18. LLM-AutoDP: Automatic Data Processing via LLM Agents for Model Fine-tuning

    cs.LG 2026-01 unverdicted novelty 6.0 of 10

    LLM agents iteratively generate and optimize data processing strategies for fine-tuning, delivering over 80% win rates versus unprocessed data and 65% versus LLM-based AutoML baselines while cutting search time by up to 10x.

  19. Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

    cs.LG 2025-07 reject novelty 6.0 of 10

    A dynamic red-teaming audit reports that 94% of MedQA-correct answers fail under adversarial mutation, with 86% privacy leak rates, 81% bias shift rates, and 66-74% hallucination rates across 15 medical LLMs.

  20. HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs

    cs.CL 2025-07 conditional novelty 6.0 of 10

    State-of-the-art LLMs are frequently inaccurate, and sometimes dangerous, when answering harm reduction questions about drug use, even when given retrieved source material.

  21. DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DynamiCare is a multi-agent LLM framework that runs multi-round diagnostic dialogues with a dynamically adjusted specialist team, evaluated on a new 500-patient benchmark built from MIMIC-III.

  22. Automating Exploratory Multiomics Research via Language Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    An LLM-based system called PROTEUS automatically explores clinical multiomics datasets and generates 360 data-driven hypotheses without human intervention, with mixed but mostly supportive external validation.

  23. HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    HSCR uses visual token dropout and logit contrast to construct self-generated dispreferred answers, then trains a medical VLM with explicit and implicit preference losses, improving zero-shot Rad-VQA, SLAKE, and PathV...

  24. Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Structured output formatting and cross-dataset fine-tuning substantially improve LLM-judge agreement with human annotations in biomedical relation extraction, though zero-shot judges are unreliable.

  25. BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum

    cs.CL 2025-05 conditional novelty 6.0 of 10

    BehaviorSFT uses `reactive` and `proactive` control tokens to fine-tune clinical LLM agents, improving their scores on the authors' new BehaviorBench dataset, though the benchmark is AI-generated and minimally clinici...

  26. PerMedCQA: Benchmarking Large Language Models on Medical Consumer Question Answering in Persian Language

    cs.CL 2025-05 reject novelty 6.0 of 10

    A first Persian consumer medical QA benchmark is released, but the paper contains no evaluation results.

  27. Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models

    cs.CV 2026-08 conditional novelty 5.0 of 10

    LoFi, trained on 4.48M medical image-text-box triplets with grounding and grounded captioning, outperforms prior models on phrase grounding, VQA, and robustness.

  28. When Language Models Meet NeuroGraphs: Exploring Enhanced Agentic LLM Framework Towards Brain Network Analysis

    cs.MA 2026-07 reject novelty 5.0 of 10

    BrainAgent, a training-free agentic LLM framework with graph understanding, knowledge retrieval, case retrieval, and reflection, claims improved but still moderate connectome classification and interpretability.

  29. Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Harrison.Rad 1.5 is a radiology-specific multimodal LLM that passes simulated FRCR 2B Short Case examinations and outperforms general-purpose frontier models on plain-film radiography reporting tasks.

  30. FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

    cs.CL 2026-07 unverdicted novelty 5.0 of 10

    FaithMed applies reinforcement learning with process-level rewards derived from evidence-based medicine rubrics to improve both task performance and reasoning faithfulness in medical LLMs.

  31. Aloe-Vision: Robust Vision-Language Models for Healthcare

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Releases open medical LVLMs trained on a quality-filtered multimodal dataset, introduces CareQA-Vision benchmark from exams, reports performance gains over baselines, and flags adversarial vulnerabilities.

  32. WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    WEQA proposes a query-adaptive agent framework combining LLMs with wearable data tools, achieving 24% higher accuracy than baselines on a benchmark from four open datasets, with gains in expert-rated usefulness.

  33. ChatHealthAI: Aligning Electronic Health Record Representations with Large Language Models for Grounded Clinical Reasoning

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    ChatHealthAI is a multimodal framework that aligns EHR foundation model representations with LLMs via a task-aware resampler for grounded clinical reasoning on longitudinal patient data while preserving predictive per...

  34. When Prompts Mislead: Textual Dominance and Diagnostic Bias in MLLMs

    cs.HC 2026-05 conditional novelty 5.0 of 10

    Textual prompts override visual cues in a medical MLLM, reducing accuracy from 75% to 46% on a hemorrhage versus drusen task even when visual grounding is present.

  35. Measuring the metacognition of AI

    cs.AI 2026-03 unverdicted novelty 5.0 of 10

    Meta-d' and signal detection theory provide quantitative tools to assess metacognitive sensitivity and risk-based regulation in large language models.

  36. Benchmarking Foundation Models with Multimodal Public Electronic Health Records

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A standardized MIMIC-IV benchmark comparing eight unimodal and multimodal foundation models shows multimodal inputs improve predictive performance without adding bias, while medical LVLMs underperform on length-of-sta...

  37. RadEyeVideo: Enhancing general-domain Large Vision Language Model for chest X-ray analysis with video representations of eye gaze

    cs.CV 2025-07 reject novelty 5.0 of 10

    A video-based eye-gaze prompt improved report generation and diagnosis for one general-purpose vision-language model, LLaVA-OneVision, but hurt or barely helped two others, and the main comparison to medical models re...

  38. Accelerating Scientific Research Through a Multi-LLM Framework

    physics.app-ph 2025-02 conditional novelty 5.0 of 10

    ARIA autonomously searches, screens, and synthesizes literature into research procedures, demonstrated on dropwise condensation.

  39. NVILA: Efficient Frontier Visual Language Models

    cs.CV 2024-12 unverdicted novelty 5.0 of 10

    NVILA improves on VILA with a scale-then-compress visual token strategy and full-lifecycle efficiency optimizations, matching or exceeding leading VLMs on image and video benchmarks while reducing training cost 1.9-5....

  40. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5 of 10

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

  41. Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA

    cs.CV 2026-07 conditional novelty 4.0 of 10

    Direct answer-only supervised fine-tuning is the most robust adaptation family on MedFrameQA, beating frozen baselines by ~6 points and outperforming complex variants on seed stability.

  42. Analyzing and Improving Fine-grained Preference Optimization in Medical LVLMs

    cs.CV 2026-06 unverdicted novelty 4.0 of 10

    Proposes bidirectional token-wise KL regularizer and visual-contrastive grounding objective to create fine-grained on-policy preference pairs for medical LVLMs by minimally editing model outputs.

  43. Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

    cs.AI 2026-05 unverdicted novelty 4.0 of 10

    HetMedAgent is a heterogeneous multi-agent framework that fuses generalist LLMs and specialist models via conflict-aware fusion and uncertainty triggers, outperforming either alone on three clinical tasks.

  44. MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs

    cs.CL 2026-02 unverdicted novelty 4.0 of 10

    MedXIAOHE is a medical MLLM that claims state-of-the-art benchmark performance through specialized pretraining to cover long-tail diseases and RL-based reasoning training.

  45. Holistic Artificial Intelligence in Medicine; improved performance and explainability

    cs.AI 2025-06 conditional novelty 4.0 of 10

    An extension of the HAIM multimodal framework that uses LLM-based retrieval and summarization to improve clinical prediction AUC from 79.9% to 90.3% and to generate document-grounded explanations.

  46. Continual Learning for Generative AI: From LLMs to MLLMs and Beyond

    cs.LG 2025-06 conditional novelty 4.0 of 10

    A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.

  47. Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Prompting multimodal LLMs with ground-truth bounding boxes and gaze durations improves chest X-ray report metrics, but the effect is inconsistent and relies on privileged annotations.

  48. Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Medical QA accuracy improves substantially through reinforcement learning with a binary correct-answer reward alone, without supervised fine-tuning on distilled reasoning traces.

  49. VLM Models and Automated Grading of Atopic Dermatitis

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Seven vision-language models were benchmarked on EASI grading of atopic dermatitis images; GPT-4o performed best, but the results rest on only five test images.

  50. QM-ToT: A Medical Tree of Thoughts Reasoning Framework for Quantized Model

    cs.CL 2025-04 unverdicted novelty 4.0 of 10

    QM-ToT applies Tree of Thoughts decomposition and evaluator layers to quantized LLMs, reporting accuracy gains from 34% to 50% on MedQAUSMLE for LLaMA2-70b and from 58.77% to 69.49% for LLaMA-3.1-8b, plus an 86.27% im...

  51. Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module

    cs.CV 2025-03 unverdicted novelty 3.0 of 10

    CA-TriNet combines co-attention transformers with a triple-LSTM module for medical report generation and reports outperforming prior models on three public datasets.

  52. MedBioLM: Optimizing Medical and Biological QA with Fine-Tuned Large Language Models and Retrieval-Augmented Generation

    cs.CL 2025-02 reject novelty 3.0 of 10

    Fine-tuning GPT-4o on biomedical QA datasets improves accuracy on MedQA, PubMedQA, and BioASQ, while RAG adds little once fine-tuning is applied.

  53. Data-Centric Foundation Models in Computational Healthcare: A Survey

    cs.LG 2024-01 unverdicted novelty 3.0 of 10

    The paper surveys data-centric strategies for foundation models in computational healthcare and supplies a curated list of related models and datasets.

Reference graph

Works this paper leans on

269 extracted references · 269 canonical work pages · cited by 53 Pith papers

  1. [1]

    M. D. Abr \`a moff, M. E. Tarver, N. Loyo-Berrios, S. Trujillo, D. Char, Z. Obermeyer, M. B. Eydelman, F. P. of Ophthalmic Imaging, D. Algorithmic Interpretation Working Group of the Collaborative Community for Ophthalmic Imaging Foundation, Washington, and W. H. Maisel. Considerations for addressing bias in artificial intelligence for health equity. NPJ ...

  2. [2]

    GPT-4 Technical Report

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. GPT-4 technical report. arXiv preprint arXiv:2303.08774, 2023

  3. [3]

    Alayrac, J

    J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35: 0 23716--23736, 2022

  4. [4]

    R. Anil, A. M. Dai, O. Firat, M. Johnson, D. Lepikhin, A. Passos, S. Shakeri, E. Taropa, P. Bailey, Z. Chen, et al. PaLM 2 technical report. arXiv preprint arXiv:2305.10403, 2023

  5. [5]

    Antaki, D

    F. Antaki, D. Milad, M. A. Chia, C.- \'E . Gigu \`e re, S. Touma, J. El-Khoury, P. A. Keane, and R. Duval. Capabilities of GPT-4 in ophthalmology: an analysis of model entropy and progress towards human-level medical question answering. British Journal of Ophthalmology, 2023

  6. [6]

    Azizi, L

    S. Azizi, L. Culp, J. Freyberg, B. Mustafa, S. Baur, S. Kornblith, T. Chen, N. Tomasev, J. Mitrovi \'c , P. Strachan, et al. Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging. Nature Biomedical Engineering, 7 0 (6): 0 756--779, 2023

  7. [7]

    Barham, A

    P. Barham, A. Chowdhery, J. Dean, S. Ghemawat, S. Hand, D. Hurt, M. Isard, H. Lim, R. Pang, S. Roy, et al. Pathways: Asynchronous distributed dataflow for ML . Proceedings of Machine Learning and Systems, 4: 0 430--449, 2022

  8. [8]

    Besta, N

    M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler. Graph of thoughts: Solving elaborate problems with large language models, 2024

Show all 269 references
  1. [9]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33: 0 1877--1901, 2020

  2. [11]

    \'A . A. Cabrera, W. Epperson, F. Hohman, M. Kahng, J. Morgenstern, and D. H. Chau. Fairvis: Visual analytics for discovering intersectional bias in machine learning. In 2019 IEEE Conference on Visual Analytics Science and Technology (VAST), pages 46--56. IEEE, 2019

  3. [13]

    D. S. Char, N. H. Shah, and D. Magnus. Implementing machine learning in health care—addressing ethical challenges. The New England journal of medicine, 378 0 (11): 0 981, 2018

  4. [14]

    W. Chen, J. Feng, J. Lu, and J. Zhou. Endo3d: Online workflow analysis for endoscopic surgeries based on 3d cnn and lstm. In OR 2.0 Context-Aware Operating Theaters, Computer Assisted Robotic Endoscopy, Clinical Image-Based Procedures, and Skin Image Analysis: First Internatio...

  5. [15]

    X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al. PaLI : A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022

  6. [16]

    Chowdhery, S

    A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. PaLM : Scaling language modeling with pathways. Journal of Machine Learning Research, 24 0 (240): 0 1--113, 2023

  7. [17]

    Cirillo, S

    D. Cirillo, S. Catuara-Solarz, C. Morey, E. Guney, L. Subirats, S. Mellino, A. Gigante, A. Valencia, M. J. Rementeria, A. S. Chadha, et al. Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare. NPJ digital medicine, 3 0 (1): 0 1--11, 2020

  8. [18]

    Claussnitzer, S

    M. Claussnitzer, S. N. Dankel, K.-H. Kim, G. Quon, W. Meuleman, C. Haugen, V. Glunk, I. S. Sousa, J. L. Beaudry, V. Puviindran, et al. Fto obesity variant circuitry and adipocyte browning in humans. New England Journal of Medicine, 373 0 (10): 0 895--907, 2015

  9. [19]

    Standards for reporting plain language summaries (pls) for cochrane diagnostic test accuracy reviews, 2014

    Cochrane. Standards for reporting plain language summaries (pls) for cochrane diagnostic test accuracy reviews, 2014. https://methods.cochrane.org/sites/methods.cochrane.org.sdt/files/uploads/Draft PLS document.pdf

  10. [20]

    E. D. Cubuk, B. Zoph, J. Shlens, and Q. V. Le. Randaugment: Practical automated data augmentation with a reduced search space. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 702--703, 2020

  11. [22]

    Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov. Transformer- XL : Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019

  12. [23]

    P. Densen. Challenges and opportunities facing medical education. Transactions of the American Clinical and Climatological Association, 122: 0 48, 2011

  13. [24]

    Devaraj, I

    A. Devaraj, I. Marshall, B. Wallace, and J. J. Li. Paragraph-level simplification of medical texts. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics, pages 4972--4984. Association for Computational Linguistics...

  14. [25]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. BERT : Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018

  15. [26]

    Driess, F

    D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, B. Ichter, A. Wahid, J. Tompson, Q. Vuong, T. Yu, et al. PaLM-E : An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023

  16. [27]

    A. V. Eriksen, S. M \"o ller, and J. Ryg. Use of GPT-4 to diagnose complex clinical cases, 2023

  17. [28]

    Feder, I

    A. Feder, I. Laish, S. Agarwal, U. Lerner, A. Atias, C. Cheung, P. Clardy, A. Peled-Cohen, R. Fellinger, H. Liu, et al. Building a clinically-focused problem list from medical notes. In Proceedings of the 13th International Workshop on Health Text Mining and Information Analys...

  18. [29]

    Fedus, B

    W. Fedus, B. Zoph, and N. Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23 0 (120): 0 1--39, 2022

  19. [31]

    E. Ford, J. A. Carroll, H. E. Smith, D. Scott, and J. A. Cassell. Extracting information from the text of electronic medical records to improve case detection: a systematic review. Journal of the American Medical Informatics Association, 23 0 (5): 0 1007--1015, 2016

  20. [33]

    Ganapathi, J

    S. Ganapathi, J. Palmer, J. E. Alderman, M. Calvert, C. Espinoza, J. Gath, M. Ghassemi, K. Heller, F. Mckay, A. Karthikesalingam, et al. Tackling bias in ai health datasets through the standing together initiative. Nature Medicine, 28 0 (11): 0 2232--2233, 2022

  21. [34]

    Y. Gao, Y. Xiong, X. Gao, K. Jia, J. Pan, Y. Bi, Y. Dai, J. Sun, Q. Guo, M. Wang, and H. Wang. Retrieval-augmented generation for large language models: A survey, 2024

  22. [36]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

    Gemini Team, Google . Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. 2024. URL https://storage.googleapis.com/deepmind-media/gemini/gemini_v1_5_report.pdf

  23. [37]

    J. W. Gichoya, I. Banerjee, A. R. Bhimireddy, J. L. Burns, L. A. Celi, L.-C. Chen, R. Correa, N. Dullerud, M. Ghassemi, S.-C. Huang, et al. Ai recognition of patient race in medical imaging: a modelling study. The Lancet Digital Health, 4 0 (6): 0 e406--e414, 2022

  24. [38]

    Golany, A

    T. Golany, A. Aides, D. Freedman, N. Rabani, Y. Liu, E. Rivlin, G. S. Corrado, Y. Matias, W. Khoury, H. Kashtan, et al. Artificial intelligence for phase recognition in complex laparoscopic cholecystectomy. Surgical Endoscopy, 36 0 (12): 0 9215--9223, 2022

  25. [40]

    E. D. Goodman, K. K. Patel, Y. Zhang, W. Locke, C. J. Kennedy, R. Mehrotra, S. Ren, M. Guan, O. Zohar, M. Downing, et al. Analyzing surgical technique in diverse open surgical videos with multitask machine learning. JAMA surgery, 159 0 (2): 0 185--192, 2024

  26. [41]

    K. K. Grandage, D. C. Slawson, and A. F. Shaughnessy. When less is more: a practical approach to searching for evidence-based answers. Journal of the Medical Library Association, 90 0 (3): 0 298, 2002

  27. [42]

    L. D. Gruppen. Clinical reasoning: defining it, teaching it, assessing it, studying it. Western Journal of Emergency Medicine, 18 0 (1): 0 4, 2017

  28. [44]

    Gupta, K

    D. Gupta, K. Attal, and D. Demner-Fushman. A dataset for medical instructional video classification and question answering. Scientific Data, 10 0 (1): 0 158, 2023

  29. [45]

    S. Hao, T. Liu, Z. Wang, and Z. Hu. ToolkenGPT : Augmenting frozen language models with massive tools via tool embeddings. Advances in neural information processing systems, 36, 2024

  30. [46]

    X. He, Z. Cai, W. Wei, Y. Zhang, L. Mou, E. Xing, and P. Xie. PathVQA : 30000+ questions for medical visual question answering. arXiv preprint arXiv:2010.12435, 2020

  31. [47]

    Horvitz, D

    E. Horvitz, D. Heckerman, B. N. Nathwani, and L. M. Fagan. Diagnostic strategies in the hypothesis-directed pathfinder system. pages 630--636, January 1984. URL https://www.microsoft.com/en-us/research/publication/diagnostic-strategies-hypothesis-directed-pathfinder-system/

  32. [48]

    Hou and Z

    W. Hou and Z. Ji. GeneTuring tests GPT models in genomics. BioRxiv, 2023

  33. [49]

    Huang and K

    J. Huang and K. C.-C. Chang. Towards reasoning in large language models: A survey, 2023

  34. [50]

    Huang, L

    J. Huang, L. Neill, M. Wittbrodt, D. Melnick, M. Klug, M. Thompson, J. Bailitz, T. Loftus, S. Malik, A. Phull, et al. Generative artificial intelligence for chest radiograph interpretation in the emergency department. JAMA Network Open, 6 0 (10): 0 e2336100--e2336100, 2023

  35. [52]

    Irvin, P

    J. Irvin, P. Rajpurkar, M. Ko, Y. Yu, S. Ciurea-Ilcus, C. Chute, H. Marklund, B. Haghgoo, R. Ball, K. Shpanskaya, et al. Che X pert: A large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the AAAI conference on artificial intelligence...

  36. [53]

    A. Iyer, G. Sen, and P. \"O stlin. The intersections of gender and class in health status and health care. Global public health, 3 0 (S1): 0 13--24, 2008

  37. [54]

    S. E. Jackson, R. A. Hackett, and A. Steptoe. Associations between age discrimination and health and wellbeing: cross-sectional and prospective analysis of the english longitudinal study of ageing. The Lancet Public Health, 4 0 (4): 0 e200--e208, 2019

  38. [55]

    P. B. Jensen, L. J. Jensen, and S. Brunak. Mining electronic health records: towards better research applications and clinical care. Nature Reviews Genetics, 13 0 (6): 0 395--405, 2012

  39. [56]

    D. Jin, E. Pan, N. Oufattole, W.-H. Weng, H. Fang, and P. Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams. Applied Sciences, 11 0 (14): 0 6421, 2021

  40. [57]

    Q. Jin, Y. Yang, Q. Chen, and Z. Lu. GeneGPT : Augmenting large language models with domain tools for improved access to biomedical information. Bioinformatics, 40 0 (2): 0 btae075, 2024

  41. [58]

    A. E. Johnson, T. J. Pollard, L. Shen, L.-w. H. Lehman, M. Feng, M. Ghassemi, B. Moody, P. Szolovits, L. Anthony Celi, and R. G. Mark. MIMIC-III , a freely accessible critical care database. Scientific data, 3 0 (1): 0 1--9, 2016

  42. [59]

    A. E. Johnson, T. J. Pollard, S. J. Berkowitz, N. R. Greenbaum, M. P. Lungren, C.-y. Deng, R. G. Mark, and S. Horng. MIMIC-CXR , a de-identified publicly available database of chest radiographs with free-text reports. Scientific data, 6 0 (1): 0 317, 2019 a

  43. [60]

    A. E. Johnson, T. J. Pollard, N. R. Greenbaum, M. P. Lungren, C.-y. Deng, Y. Peng, Z. Lu, R. G. Mark, S. J. Berkowitz, and S. Horng. MIMIC-CXR-JPG , a large publicly available database of labeled chest radiographs. arXiv preprint arXiv:1901.07042, 2019 b

  44. [61]

    Kanjee, B

    Z. Kanjee, B. Crowe, and A. Rodman. Accuracy of a generative artificial intelligence model in a complex diagnostic challenge. Jama, 330 0 (1): 0 78--80, 2023

  45. [62]

    J. A. Kent, V. Patel, and N. A. Varela. Gender disparities in health care. Mount Sinai Journal of Medicine: A Journal of Translational and Personalized Medicine, 79 0 (5): 0 555--559, 2012

  46. [63]

    u r Evidenz, Fortbildung und Qualit \

    I. Klerings, A. S. Weinhandl, and K. J. Thaler. Information overload in healthcare: too much of a good thing? Zeitschrift f \"u r Evidenz, Fortbildung und Qualit \"a t im Gesundheitswesen , 109 0 (4-5): 0 285--290, 2015

  47. [64]

    Kouzy, J

    R. Kouzy, J. Abi Jaoude, A. Kraitem, M. B. El Alam, B. Karam, E. Adib, J. Zarka, C. Traboulsi, E. W. Akl, and K. Baddour. Coronavirus goes viral: quantifying the covid-19 misinformation epidemic on twitter. Cureus, 12 0 (3), 2020

  48. [65]

    Laber, S

    S. Laber, S. Forcisi, L. Bentley, J. Petzold, F. Moritz, K. S. Smirnov, L. Al Sadat, I. Williamson, S. Strobel, T. Agnew, et al. Linking the fto obesity rs1421085 variant circuitry to cellular, metabolic, and organismal phenotypes in vivo. Science advances, 7 0 (30): 0 eabg0108, 2021

  49. [66]

    Le Scao, A

    T. Le Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili \'c , D. Hesslow, R. Castagn \'e , A. S. Luccioni, F. Yvon, M. Gall \'e , et al. Bloom: A 176b-parameter open-access multilingual language model. 2022

  50. [67]

    Leifman, A

    G. Leifman, A. Aides, T. Golany, D. Freedman, and E. Rivlin. Pixel-accurate segmentation of surgical tools based on bounding box annotations. In 2022 26th International Conference on Pattern Recognition (ICPR), pages 5096--5103. IEEE, 2022

  51. [69]

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao. LLaVa-Med : Training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems, 36, 2024

  52. [70]

    Y. Li, R. M. Wehbe, F. S. Ahmad, H. Wang, and Y. Luo. A comparative study of pretrained language models for long clinical text. Journal of the American Medical Informatics Association, 30 0 (2): 0 340--347, 2023

  53. [71]

    Liu, L.-M

    B. Liu, L.-M. Zhan, L. Xu, L. Ma, Y. Yang, and X.-M. Wu. Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering. In 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI), pages 1650--1654. IEEE, 2021

  54. [72]

    M. Liu, Y. Ning, S. Teixayavong, M. Mertens, J. Xu, D. S. W. Ting, L. T.-E. Cheng, J. C. L. Ong, Z. L. Teo, T. F. Tan, et al. A translational perspective towards clinical ai fairness. NPJ Digital Medicine, 6 0 (1): 0 172, 2023

  55. [73]

    N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12: 0 157--173, 2024

  56. [74]

    R. J. Loos and G. S. Yeo. The genetics of obesity: from discovery to biology. Nature Reviews Genetics, 23 0 (2): 0 120--133, 2022

  57. [75]

    L \'o pez and V

    N. L \'o pez and V. L. Gadsden. Health inequities, social determinants, and intersectionality. In Perspectives on health equity and social determinants of health. National Academies Press (US), 2017

  58. [77]

    R. Luo, L. Sun, Y. Xia, T. Qin, S. Zhang, H. Poon, and T.-Y. Liu. BioGPT : generative pre-trained transformer for biomedical text generation and mining. Briefings in bioinformatics, 23 0 (6): 0 bbac409, 2022

  59. [81]

    Medina-Mart \' nez, C

    J. Medina-Mart \' nez, C. Saus-Ortega, M. M. S \'a nchez-Lorente, E. M. Sosa-Palanca, P. Garc \' a-Mart \' nez, and M. I. M \'a rmol-L \'o pez. Health inequities in lgbt people and nursing interventions to reduce them: A systematic review. International Journal of Environmenta...

  60. [82]

    Papers with code - medical, 2024

    Meta. Papers with code - medical, 2024. URL https://paperswithcode.com/area/medical. Accessed: 2024-04-26

  61. [83]

    M. Moor, O. Banerjee, Z. S. H. Abad, H. M. Krumholz, J. Leskovec, E. J. Topol, and P. Rajpurkar. Foundation models for generalist medical artificial intelligence. Nature, 616 0 (7956): 0 259--265, 2023 a

  62. [84]

    M. Moor, Q. Huang, S. Wu, M. Yasunaga, Y. Dalmia, J. Leskovec, C. Zakka, E. P. Reis, and P. Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. In Machine Learning for Health (ML4H), pages 353--367. PMLR, 2023 b

  63. [85]

    Nakano, J

    R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V. Kosaraju, W. Saunders, et al. WebGPT : Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332, 2021

  64. [87]

    Novin and E

    A. Novin and E. Meyers. Making sense of conflicting science information: Exploring bias in the search engine result page. In Proceedings of the 2017 conference on conference human information interaction and retrieval, pages 175--184, 2017

  65. [88]

    C. I. Nwoye, D. Mutter, J. Marescaux, and N. Padoy. Weakly supervised convolutional lstm approach for tool tracking in laparoscopic videos. International journal of computer assisted radiology and surgery, 14: 0 1059--1067, 2019

  66. [89]

    Obermeyer, B

    Z. Obermeyer, B. Powers, C. Vogeli, and S. Mullainathan. Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366 0 (6464): 0 447--453, 2019

  67. [90]

    J. Oh, G. Lee, S. Bae, J.-m. Kwon, and E. Choi. Ecg-qa: A comprehensive question answering dataset combined with electrocardiogram. In A. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pa...

  68. [91]

    J. A. Omiye, J. C. Lester, S. Spichak, V. Rotemberg, and R. Daneshjou. Large language models propagate race-based medicine. NPJ Digital Medicine, 6 0 (1): 0 195, 2023

  69. [92]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35: 0 27730--27744, 2022

  70. [93]

    A. G. Pacheco, G. R. Lima, A. S. Salomao, B. Krohling, I. P. Biral, G. G. de Angelo, F. C. Alves Jr, J. G. Esgario, A. C. Simora, P. B. Castro, et al. PAD-UFES-20 : A skin lesion dataset composed of patient data and clinical images collected from smartphones. Data in brief, 32...

  71. [94]

    Parmar, A

    M. Parmar, A. Naik, H. Gupta, D. Agrawal, and C. Baral. LongBoX : Evaluating transformers on long-sequence clinical tasks, 2023

  72. [95]

    Pelka, S

    O. Pelka, S. Koitka, J. R \"u ckert, F. Nensa, and C. M. Friedrich. Radiology objects in context (roco): a multimodal image dataset. In Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis: 7th Joint Inte...

  73. [97]

    M. Poli, S. Massaroli, E. Nguyen, D. Y. Fu, T. Dao, S. Baccus, Y. Bengio, S. Ermon, and C. R \'e . Hyena hierarchy: Towards larger convolutional language models. In International Conference on Machine Learning, pages 28043--28078. PMLR, 2023

  74. [98]

    S. Qiao, Y. Ou, N. Zhang, X. Chen, Y. Yao, S. Deng, C. Tan, F. Huang, and H. Chen. Reasoning with language model prompting: A survey, 2023

  75. [99]

    Y. Qin, S. Liang, Y. Ye, K. Zhu, L. Yan, Y. Lu, Y. Lin, X. Cong, X. Tang, B. Qian, et al. ToolLLM : Facilitating large language models to master 16000+ real-world apis. arXiv preprint arXiv:2307.16789, 2023

  76. [100]

    Radford, K

    A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al. Improving language understanding by generative pre-training. 2018

  77. [101]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21 0 (1): 0 5485--5551, 2020

  78. [102]

    Rajpurkar, E

    P. Rajpurkar, E. Chen, O. Banerjee, and E. J. Topol. AI in health and medicine. Nature medicine, 28 0 (1): 0 31--38, 2022

  79. [103]

    Ramesh, N

    V. Ramesh, N. A. Chi, and P. Rajpurkar. Improving radiology report generation systems by removing hallucinated references to non-existent priors. In A. Parziale, M. Agrawal, S. Joshi, I. Y. Chen, S. Tang, L. Oala, and A. Subbaswamy, editors, Proceedings of the 2nd Machine Lear...

  80. [104]

    M. S. Razai, H. K. Kankam, A. Majeed, A. Esmail, and D. R. Williams. Mitigating ethnic disparities in covid-19 and beyond. bmj, 372, 2021

  81. [105]

    M. S. R \' os, M. A. Molina-Rodriguez, D. Londo \ n o, C. A. Guill \'e n, S. Sierra, F. Zapata, and L. F. Giraldo. Cholec80-cvs: An open dataset with an evaluation of strasberg’s critical view of safety for ai. Scientific Data, 10 0 (1): 0 194, 2023

  82. [106]

    D. E. Sanford and S. M. Strasberg. A simple effective method for generation of a permanent record of the critical view of safety during laparoscopic cholecystectomy by intraoperative “doublet” photography. Journal of the American College of Surgeons, 218 0 (2): 0 170--178, 2014

  83. [107]

    Sbaffi, J

    L. Sbaffi, J. Walton, J. Blenkinsopp, and G. Walton. Information overload in emergency medicine physicians: a multisite case study exploring the causes, impact, and solutions in four north england national health service trusts. Journal of medical Internet research, 22 0 (7): ...

  84. [108]

    Schick, J

    T. Schick, J. Dwivedi-Yu, R. Dess \` , R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 2024

  85. [110]

    Sieferd, N

    E. Sieferd, N. Mohanty, and R. J. Holden. After visit summary: Not an afterthought. In Proceedings of the International Symposium on Human Factors and Ergonomics in Health Care, volume 8, pages 85--89. SAGE Publications Sage CA: Los Angeles, CA, 2019

  86. [111]

    Singhal, S

    K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al. Large language models encode clinical knowledge. Nature, 620 0 (7972): 0 172--180, 2023 a

  87. [114]

    Steptoe and P

    A. Steptoe and P. Zaninotto. Lower socioeconomic status and the acceleration of aging: An outcome-wide analysis. Proceedings of the National Academy of Sciences, 117 0 (26): 0 14911--14917, 2020

  88. [115]

    S. M. Strasberg and M. L. Brunt. Rationale and use of the critical view of safety in laparoscopic cholecystectomy. Journal of the American College of Surgeons, 211 0 (1): 0 132--138, 2010

  89. [116]

    Stutz, A

    D. Stutz, A. T. Cemgil, A. G. Roy, T. Matejovicova, M. Barsbey, P. Strachan, M. Schaekermann, J. Freyberg, R. Rikhye, B. Freeman, J. P. Matos, U. Telang, D. R. Webster, Y. Liu, G. S. Corrado, Y. Matias, P. Kohli, Y. Liu, A. Doucet, and A. Karthikesalingam. Evaluating AI system...

  90. [117]

    Tanno, D

    R. Tanno, D. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, K. Singhal, S. Azizi, T. Tu, et al. Consensus, dissensus and synergy between clinicians and specialist foundation models in radiology report generation. 2024

  91. [118]

    Image challenge

    The New England Journal of Medicine . Image challenge. https://www.nejm.org/image-challenge, 2024

  92. [121]

    T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P.-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena, et al. Towards generalist biomedical AI . NEJM AI, 1 0 (3): 0 AIoa2300138, 2024 a

  93. [122]

    T. Tu, A. Palepu, M. Schaekermann, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, N. Tomasev, et al. Towards conversational diagnostic AI . arXiv preprint arXiv:2401.05654, 2024 b

  94. [123]

    A. P. Twinanda, S. Shehata, D. Mutter, J. Marescaux, M. De Mathelin, and N. Padoy. Endonet: a deep architecture for recognition tasks on laparoscopic videos. IEEE transactions on medical imaging, 36 0 (1): 0 86--97, 2016

  95. [126]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, . Kaiser, and I. Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017

  96. [127]

    Wagner, N

    P. Wagner, N. Strodthoff, R.-D. Bousseljot, D. Kreiseler, F. I. Lunze, W. Samek, and T. Schaeffter. PTB-XL , a large publicly available electrocardiography dataset. Scientific data, 7 0 (1): 0 1--15, 2020

  97. [128]

    Z. Wan, C. Liu, X. Wang, C. Tao, H. Shen, Z. Peng, J. Fu, R. Arcucci, H. Yao, and M. Zhang. Electrocardiogram instruction tuning for report generation, 2024

  98. [129]

    A. Wang, V. V. Ramaswamy, and O. Russakovsky. Towards intersectionality in machine learning: Including more identities, handling underrepresentation, and performing evaluation. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 336--...

  99. [131]

    Y. Wang, W. Chen, X. Han, X. Lin, H. Zhao, Y. Liu, B. Zhai, J. Yuan, Q. You, and H. Yang. Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning, 2024

  100. [133]

    L. W. Way, L. Stewart, W. Gantert, K. Liu, C. M. Lee, K. Whang, and J. G. Hunter. Causes and prevention of laparoscopic bile duct injuries: analysis of 252 cases from a human factors and cognitive psychology perspective. Annals of surgery, 237 0 (4): 0 460--469, 2003

  101. [135]

    J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35: 0 24824--24837, 2022

  102. [136]

    Weng and B

    Y. Weng and B. Li. Visual answer localization with cross-modal mutual knowledge transfer. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1--5. IEEE, 2023

  103. [137]

    D. R. Williams and R. Wyatt. Racial bias in health care and health: challenges and opportunities. Jama, 314 0 (6): 0 555--556, 2015

  104. [140]

    F. Yang, M. Cisse, and S. Koyejo. Fairness with overlapping groups; a probabilistic perspective. Advances in neural information processing systems, 33: 0 4067--4078, 2020

  105. [141]

    S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan. Tree of thoughts: Deliberate problem solving with large language models, 2023

  106. [143]

    X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, et al. MMMU : A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi. arXiv preprint arXiv:2311.16502, 2023

  107. [144]

    Zakka, R

    C. Zakka, R. Shad, A. Chaurasia, A. R. Dalal, J. L. Kim, M. Moor, R. Fong, C. Phillips, K. Alexander, E. Ashley, et al. Almanac—retrieval-augmented language models for clinical medicine. NEJM AI, 1 0 (2): 0 AIoa2300068, 2024

  108. [146]

    Zelikman, J

    E. Zelikman, J. Mu, N. D. Goodman, and Y. T. Wu. Star: Self-taught reasoner bootstrapping reasoning with reasoning. Advances in Neural Information Processing Systems (NeurIPS), 2022

  109. [148]

    D. Zhou, N. Schärli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, C. Cui, O. Bousquet, Q. Le, and E. Chi. Least-to-most prompting enables complex reasoning in large language models, 2023

  110. [149]

    arXiv preprint arXiv:2312.11805 , year=

    Gemini: A family of highly capable multimodal models , author=. arXiv preprint arXiv:2312.11805 , year=

  111. [150]

    Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context , author=

  112. [151]

    2016 , publisher=

    Johnson, Alistair EW and Pollard, Tom J and Shen, Lu and Lehman, Li-wei H and Feng, Mengling and Ghassemi, Mohammad and Moody, Benjamin and Szolovits, Peter and Anthony Celi, Leo and Mark, Roger G , journal=. 2016 , publisher=

  113. [152]

    Applied Sciences , volume=

    What disease does this patient have? a large-scale open domain question answering dataset from medical exams , author=. Applied Sciences , volume=. 2021 , publisher=

  114. [153]

    arXiv preprint arXiv:2311.13668 , year=

    Hyland, Stephanie L and Bannur, Shruthi and Bouzid, Kenza and Castro, Daniel C and Ranjit, Mercy and Schwaighofer, Anton and P. arXiv preprint arXiv:2311.13668 , year=

  115. [154]

    Proceedings of the 2nd Machine Learning for Health symposium , pages =

    Improving Radiology Report Generation Systems by Removing Hallucinated References to Non-existent Priors , author =. Proceedings of the 2nd Machine Learning for Health symposium , pages =. 2022 , editor =

  116. [155]

    Irvin, Jeremy and Rajpurkar, Pranav and Ko, Michael and Yu, Yifan and Ciurea-Ilcus, Silviana and Chute, Chris and Marklund, Henrik and Haghgoo, Behzad and Ball, Robyn and Shpanskaya, Katie and others , booktitle=. Che

  117. [156]

    Consensus, dissensus and synergy between clinicians and specialist foundation models in radiology report generation , author=

  118. [157]

    Scientific data , volume=

    A dataset of clinically generated visual questions and answers about radiology images , author=. Scientific data , volume=. 2018 , publisher=

  119. [158]

    arXiv preprint arXiv:2402.18545 , year=

    Crowdsourcing Dermatology Images with Google Search Ads: Creating a Real-World Skin Condition Dataset , author=. arXiv preprint arXiv:2402.18545 , year=

  120. [159]

    2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI) , pages=

    Slake: A semantically-labeled knowledge-enhanced dataset for medical visual question answering , author=. 2021 IEEE 18th International Symposium on Biomedical Imaging (ISBI) , pages=. 2021 , organization=

  121. [160]

    Radiology Objects in COntext (ROCO): a multimodal image dataset , author=. Intravascular Imaging and Computer Assisted Stenting and Large-Scale Annotation of Biomedical Data and Expert Label Synthesis: 7th Joint International Workshop, CVII-STENT 2018 and Third International W...

  122. [161]

    ECG-QA: A Comprehensive Question Answering Dataset Combined With Electrocardiogram , url =

    Oh, Jungwoo and Lee, Gyubok and Bae, Seongsu and Kwon, Joon-myoung and Choi, Edward , booktitle =. ECG-QA: A Comprehensive Question Answering Dataset Combined With Electrocardiogram , url =

  123. [162]

    and Fagan, Lawrence M

    Horvitz, Eric and Heckerman, David and Nathwani, Bharat N. and Fagan, Lawrence M. , title =. 1984 , month =

  124. [163]

    JAMA Network Open , volume=

    Generative Artificial Intelligence for Chest Radiograph Interpretation in the Emergency Department , author=. JAMA Network Open , volume=. 2023 , publisher=

  125. [164]

    He, Xuehai and Cai, Zhuo and Wei, Wenlan and Zhang, Yichen and Mou, Luntian and Xing, Eric and Xie, Pengtao , journal=

  126. [165]

    2019 , publisher=

    Johnson, Alistair EW and Pollard, Tom J and Berkowitz, Seth J and Greenbaum, Nathaniel R and Lungren, Matthew P and Deng, Chih-ying and Mark, Roger G and Horng, Steven , journal=. 2019 , publisher=

  127. [166]

    Johnson, Alistair EW and Pollard, Tom J and Greenbaum, Nathaniel R and Lungren, Matthew P and Deng, Chih-ying and Peng, Yifan and Lu, Zhiyong and Mark, Roger G and Berkowitz, Seth J and Horng, Steven , journal=

  128. [167]

    Data in brief , volume=

    Pacheco, Andre GC and Lima, Gustavo R and Salomao, Amanda S and Krohling, Breno and Biral, Igor P and de Angelo, Gabriel G and Alves Jr, F. Data in brief , volume=. 2020 , publisher=

  129. [168]

    Scientific Data , volume=

    A dataset for medical instructional video classification and question answering , author=. Scientific Data , volume=. 2023 , publisher=

  130. [169]

    arXiv preprint arXiv:2112.07219 , year=

    A real-time spatiotemporal AI model analyzes skill in open surgical videos , author=. arXiv preprint arXiv:2112.07219 , year=

  131. [170]

    JAMA surgery , volume=

    Analyzing surgical technique in diverse open surgical videos with multitask machine learning , author=. JAMA surgery , volume=. 2024 , publisher=

  132. [171]

    2023 , publisher=

    Hou, Wenpin and Ji, Zhicheng , journal=. 2023 , publisher=

  133. [172]

    2024 , publisher=

    Jin, Qiao and Yang, Yifan and Chen, Qingyu and Lu, Zhiyong , journal=. 2024 , publisher=

  134. [173]

    2022 , publisher=

    Luo, Renqian and Sun, Liai and Xia, Yingce and Qin, Tao and Zhang, Sheng and Poon, Hoifung and Liu, Tie-Yan , journal=. 2022 , publisher=

  135. [174]

    Advances in neural information processing systems , volume=

    Language models are few-shot learners , author=. Advances in neural information processing systems , volume=

  136. [175]

    Towards generalist biomedical

    Tu, Tao and Azizi, Shekoofeh and Driess, Danny and Schaekermann, Mike and Amin, Mohamed and Chang, Pi-Chuan and Carroll, Andrew and Lau, Charles and Tanno, Ryutaro and Ktena, Ira and others , journal=. Towards generalist biomedical. 2024 , publisher=

  137. [176]

    arXiv preprint arXiv:2205.01917 , year=

    Coca: Contrastive captioners are image-text foundation models , author=. arXiv preprint arXiv:2205.01917 , year=

  138. [177]

    Advances in neural information processing systems , volume=

    Flamingo: a visual language model for few-shot learning , author=. Advances in neural information processing systems , volume=

  139. [178]

    arXiv preprint arXiv:2312.00164 , year=

    Towards accurate differential diagnosis with large language models , author=. arXiv preprint arXiv:2312.00164 , year=

  140. [179]

    arXiv preprint arXiv:2311.16452 , year=

    Can generalist foundation models outcompete special-purpose tuning? case study in medicine , author=. arXiv preprint arXiv:2311.16452 , year=

  141. [180]

    ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

    Learning to locate visual answer in video corpus using question , author=. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=. 2023 , organization=

  142. [181]

    arXiv preprint arXiv:2203.06667 , year=

    Towards visual-prompt temporal answering grounding in medical instructional video , author=. arXiv preprint arXiv:2203.06667 , year=

  143. [182]

    arXiv preprint arXiv:2311.05591 , year=

    Accuracy of a vision-language model on challenging medical cases , author=. arXiv preprint arXiv:2311.05591 , year=

  144. [183]

    Performance of multimodal

    Yang, Zhichao and Yao, Zonghai and Tasmin, Mahbuba and Vashisht, Parth and Jang, Won Seok and Ouyang, Feiyun and Wang, Beining and Berlowitz, Dan and Yu, Hong , journal=. Performance of multimodal. 2023 , publisher=

  145. [184]

    Proceedings of the 13th International Workshop on Health Text Mining and Information Analysis (LOUHI) , pages=

    Building a clinically-focused problem list from medical notes , author=. Proceedings of the 13th International Workshop on Health Text Mining and Information Analysis (LOUHI) , pages=

  146. [185]

    Advances in neural information processing systems , volume=

    Attention is all you need , author=. Advances in neural information processing systems , volume=

  147. [186]

    Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , journal=

  148. [187]

    The Journal of Machine Learning Research , volume=

    Exploring the limits of transfer learning with a unified text-to-text transformer , author=. The Journal of Machine Learning Research , volume=. 2020 , publisher=

  149. [188]

    arXiv preprint arXiv:2109.01652 , year=

    Finetuned language models are zero-shot learners , author=. arXiv preprint arXiv:2109.01652 , year=

  150. [189]

    Anil, Rohan and Dai, Andrew M and Firat, Orhan and Johnson, Melvin and Lepikhin, Dmitry and Passos, Alexandre and Shakeri, Siamak and Taropa, Emanuel and Bailey, Paige and Chen, Zhifeng and others , journal=

  151. [190]

    Pathways: Asynchronous distributed dataflow for

    Barham, Paul and Chowdhery, Aakanksha and Dean, Jeff and Ghemawat, Sanjay and Hand, Steven and Hurt, Daniel and Isard, Michael and Lim, Hyeontaek and Pang, Ruoming and Roy, Sudip and others , journal=. Pathways: Asynchronous distributed dataflow for

  152. [191]

    Chowdhery, Aakanksha and Narang, Sharan and Devlin, Jacob and Bosma, Maarten and Mishra, Gaurav and Roberts, Adam and Barham, Paul and Chung, Hyung Won and Sutton, Charles and Gehrmann, Sebastian and others , journal=

  153. [192]

    2018 , publisher=

    Improving language understanding by generative pre-training , author=. 2018 , publisher=

  154. [193]

    arXiv preprint arXiv:2302.13971 , year=

    Touvron, Hugo and Lavril, Thibaut and Izacard, Gautier and Martinet, Xavier and Lachaux, Marie-Anne and Lacroix, Timoth. arXiv preprint arXiv:2302.13971 , year=

  155. [194]

    Nature , volume=

    Large language models encode clinical knowledge , author=. Nature , volume=. 2023 , publisher=

  156. [195]

    Towards conversational diagnostic

    Tu, Tao and Palepu, Anil and Schaekermann, Mike and Saab, Khaled and Freyberg, Jan and Tanno, Ryutaro and Wang, Amy and Li, Brenna and Amin, Mohamed and Tomasev, Nenad and others , journal=. Towards conversational diagnostic

  157. [196]

    Western Journal of Emergency Medicine , volume=

    Clinical reasoning: defining it, teaching it, assessing it, studying it , author=. Western Journal of Emergency Medicine , volume=. 2017 , publisher=

  158. [197]

    Jain, Saahil and Agrawal, Ashwin and Saporta, Adriel and Truong, Steven QH and Duong, Du Nguyen and Bui, Tan and Chambon, Pierre and Zhang, Yuhao and Lungren, Matthew P and Ng, Andrew Y and others , journal=

  159. [198]

    arXiv preprint arXiv:2308.01834 , year=

    The Capability of Large Language Models to Measure Psychiatric Functioning , author=. arXiv preprint arXiv:2308.01834 , year=

  160. [199]

    NPJ Digital Medicine , volume=

    A translational perspective towards clinical AI fairness , author=. NPJ Digital Medicine , volume=. 2023 , publisher=

  161. [200]

    The Lancet Digital Health , volume=

    AI recognition of patient race in medical imaging: a modelling study , author=. The Lancet Digital Health , volume=. 2022 , publisher=

  162. [201]

    Science , volume=

    Dissecting racial bias in an algorithm used to manage the health of populations , author=. Science , volume=. 2019 , publisher=

  163. [202]

    Advances in neural information processing systems , volume=

    Fairness with overlapping groups; a probabilistic perspective , author=. Advances in neural information processing systems , volume=

  164. [203]

    arXiv preprint arXiv:2104.08666 , year=

    Worst of both worlds: Biases compound in pre-trained vision-and-language models , author=. arXiv preprint arXiv:2104.08666 , year=

  165. [204]

    arXiv preprint arXiv:2304.13855 , year=

    Multimodal composite association score: Measuring gender bias in generative multimodal models , author=. arXiv preprint arXiv:2304.13855 , year=

  166. [205]

    Proceedings of the 2017 conference on conference human information interaction and retrieval , pages=

    Making sense of conflicting science information: Exploring bias in the search engine result page , author=. Proceedings of the 2017 conference on conference human information interaction and retrieval , pages=

  167. [206]

    2019 IEEE Conference on Visual Analytics Science and Technology (VAST) , pages=

    FairVis: Visual analytics for discovering intersectional bias in machine learning , author=. 2019 IEEE Conference on Visual Analytics Science and Technology (VAST) , pages=. 2019 , organization=

  168. [207]

    NPJ digital medicine , volume=

    Sex and gender differences and biases in artificial intelligence for biomedicine and healthcare , author=. NPJ digital medicine , volume=. 2020 , publisher=

  169. [208]

    Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

    Towards intersectionality in machine learning: Including more identities, handling underrepresentation, and performing evaluation , author=. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

  170. [209]

    Perspectives on health equity and social determinants of health , year=

    Health inequities, social determinants, and intersectionality , author=. Perspectives on health equity and social determinants of health , year=

  171. [210]

    The Lancet Public Health , volume=

    Associations between age discrimination and health and wellbeing: cross-sectional and prospective analysis of the English Longitudinal Study of Ageing , author=. The Lancet Public Health , volume=. 2019 , publisher=

  172. [211]

    International Journal of Environmental Research and Public Health , volume=

    Health inequities in LGBT people and nursing interventions to reduce them: A systematic review , author=. International Journal of Environmental Research and Public Health , volume=. 2021 , publisher=

  173. [212]

    Proceedings of the National Academy of Sciences , volume=

    Lower socioeconomic status and the acceleration of aging: An outcome-wide analysis , author=. Proceedings of the National Academy of Sciences , volume=. 2020 , publisher=

  174. [213]

    bmj , volume=

    Mitigating ethnic disparities in covid-19 and beyond , author=. bmj , volume=. 2021 , publisher=

  175. [214]

    Jama , volume=

    Racial bias in health care and health: challenges and opportunities , author=. Jama , volume=. 2015 , publisher=

  176. [215]

    Global public health , volume=

    The intersections of gender and class in health status and health care , author=. Global public health , volume=. 2008 , publisher=

  177. [216]

    Mount Sinai Journal of Medicine: A Journal of Translational and Personalized Medicine , volume=

    Gender disparities in health care , author=. Mount Sinai Journal of Medicine: A Journal of Translational and Personalized Medicine , volume=. 2012 , publisher=

  178. [217]

    The New England journal of medicine , volume=

    Implementing machine learning in health care—addressing ethical challenges , author=. The New England journal of medicine , volume=. 2018 , publisher=

  179. [218]

    Nature Medicine , volume=

    Tackling bias in AI health datasets through the STANDING Together initiative , author=. Nature Medicine , volume=. 2022 , publisher=

  180. [219]

    2023 , organization=

    Khanna, Sameer and Dejl, Adam and Yoon, Kibo and Truong, Steven QH and Duong, Hanh and Saenz, Agustina and Rajpurkar, Pranav , booktitle=. 2023 , organization=

  181. [220]

    Webster and Yuan Liu and Greg S

    David Stutz and Ali Taylan Cemgil and Abhijit Guha Roy and Tatiana Matejovicova and Melih Barsbey and Patricia Strachan and Mike Schaekermann and Jan Freyberg and Rajeev Rikhye and Beverly Freeman and Javier Perez Matos and Umesh Telang and Dale R. Webster and Yuan Liu and Gre...

  182. [221]

    2023 , eprint=

    Reasoning with Language Model Prompting: A Survey , author=. 2023 , eprint=

  183. [222]

    2023 , eprint=

    Towards Reasoning in Large Language Models: A Survey , author=. 2023 , eprint=

  184. [223]

    2024 , eprint=

    Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning , author=. 2024 , eprint=

  185. [224]

    Advances in neural information processing systems , volume=

    Chain-of-thought prompting elicits reasoning in large language models , author=. Advances in neural information processing systems , volume=

  186. [225]

    2023 , eprint=

    Least-to-Most Prompting Enables Complex Reasoning in Large Language Models , author=. 2023 , eprint=

  187. [226]

    2023 , eprint=

    Tree of Thoughts: Deliberate Problem Solving with Large Language Models , author=. 2023 , eprint=

  188. [227]

    2024 , eprint=

    Graph of Thoughts: Solving Elaborate Problems with Large Language Models , author=. 2024 , eprint=

  189. [228]

    Advances in Neural Information Processing Systems (NeurIPS) , year=

    Star: Self-taught reasoner bootstrapping reasoning with reasoning , author=. Advances in Neural Information Processing Systems (NeurIPS) , year=

  190. [229]

    arXiv preprint arXiv:2203.11171 , year=

    Self-consistency improves chain of thought reasoning in language models , author=. arXiv preprint arXiv:2203.11171 , year=

  191. [230]

    Advances in Neural Information Processing Systems , volume=

    Toolformer: Language models can teach themselves to use tools , author=. Advances in Neural Information Processing Systems , volume=

  192. [231]

    Hao, Shibo and Liu, Tianyang and Wang, Zhen and Hu, Zhiting , journal=

  193. [232]

    2024 , eprint=

    Retrieval-Augmented Generation for Large Language Models: A Survey , author=. 2024 , eprint=

  194. [233]

    Driess, Danny and Xia, Fei and Sajjadi, Mehdi SM and Lynch, Corey and Chowdhery, Aakanksha and Ichter, Brian and Wahid, Ayzaan and Tompson, Jonathan and Vuong, Quan and Yu, Tianhe and others , journal=

  195. [234]

    Bloom: A 176b-parameter open-access multilingual language model , author=

  196. [235]

    Chen, Xi and Wang, Xiao and Changpinyo, Soravit and Piergiovanni, AJ and Padlewski, Piotr and Salz, Daniel and Goodman, Sebastian and Grycner, Adam and Mustafa, Basil and Beyer, Lucas and others , journal=

  197. [236]

    Nature , volume=

    Foundation models for generalist medical artificial intelligence , author=. Nature , volume=. 2023 , publisher=

  198. [237]

    Qin, Yujia and Liang, Shihao and Ye, Yining and Zhu, Kunlun and Yan, Lan and Lu, Yaxi and Lin, Yankai and Cong, Xin and Tang, Xiangru and Qian, Bill and others , journal=

  199. [238]

    Nakano, Reiichiro and Hilton, Jacob and Balaji, Suchir and Wu, Jeff and Ouyang, Long and Kim, Christina and Hesse, Christopher and Jain, Shantanu and Kosaraju, Vineet and Saunders, William and others , journal=

  200. [239]

    arXiv preprint arXiv:2305.17126 , year=

    Large language models as tool makers , author=. arXiv preprint arXiv:2305.17126 , year=

  201. [240]

    Li, Chunyuan and Wong, Cliff and Zhang, Sheng and Usuyama, Naoto and Liu, Haotian and Yang, Jianwei and Naumann, Tristan and Poon, Hoifung and Gao, Jianfeng , journal=

  202. [241]

    arXiv preprint arXiv:2312.07814 , year=

    A Foundational Multimodal Vision Language AI Assistant for Human Pathology , author=. arXiv preprint arXiv:2312.07814 , year=

  203. [242]

    arXiv preprint arXiv:2403.08002 , year=

    Training Small Multimodal Models to Bridge Biomedical Competency Gap: A Case Study in Radiology Imaging , author=. arXiv preprint arXiv:2403.08002 , year=

  204. [243]

    Xu, Shawn and Yang, Lin and Kelly, Christopher and Sieniek, Marcin and Kohlberger, Timo and Ma, Martin and Weng, Wei-Hung and Kiraly, Attila and Kazemzadeh, Sahar and Melamed, Zakkai and others , journal=

  205. [244]

    Machine Learning for Healthcare Conference , pages=

    Contrastive learning of medical visual representations from paired images and text , author=. Machine Learning for Healthcare Conference , pages=. 2022 , organization=

  206. [245]

    2024 , eprint=

    Electrocardiogram Instruction Tuning for Report Generation , author=. 2024 , eprint=

  207. [246]

    Journal of the American Medical Informatics Association , volume=

    A comparative study of pretrained language models for long clinical text , author=. Journal of the American Medical Informatics Association , volume=. 2023 , publisher=

  208. [247]

    2311.09564 , archivePrefix=

    Mihir Parmar and Aakanksha Naik and Himanshu Gupta and Disha Agrawal and Chitta Baral , year=. 2311.09564 , archivePrefix=

  209. [248]

    Transactions of the Association for Computational Linguistics , volume=

    Lost in the middle: How language models use long contexts , author=. Transactions of the Association for Computational Linguistics , volume=. 2024 , publisher=

  210. [249]

    arXiv preprint arXiv:2204.06683 , year=

    Revisiting transformer-based models for long document classification , author=. arXiv preprint arXiv:2204.06683 , year=

  211. [250]

    Transformer-

    Dai, Zihang and Yang, Zhilin and Yang, Yiming and Carbonell, Jaime and Le, Quoc V and Salakhutdinov, Ruslan , journal=. Transformer-

  212. [251]

    International Conference on Machine Learning , pages=

    Hyena hierarchy: Towards larger convolutional language models , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  213. [252]

    Overview of the M ed V id QA 2022 Shared Task on Medical Video Question-Answering

    Gupta, Deepak and Demner-Fushman, Dina. Overview of the M ed V id QA 2022 Shared Task on Medical Video Question-Answering. Proceedings of the 21st Workshop on Biomedical Language Processing. 2022. doi:10.18653/v1/2022.bionlp-1.25

  214. [253]

    Paragraph-level Simplification of Medical Texts

    Devaraj, Ashwin and Marshall, Iain and Wallace, Byron and Li, Junyi Jessy. Paragraph-level Simplification of Medical Texts. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics. 2021

  215. [254]

    Standards for reporting Plain Language Summaries (PLS) for Cochrane Diagnostic Test Accuracy Reviews

    Cochrane. Standards for reporting Plain Language Summaries (PLS) for Cochrane Diagnostic Test Accuracy Reviews

  216. [255]

    Proceedings of the International Symposium on Human Factors and Ergonomics in Health Care , volume=

    After Visit Summary: Not an Afterthought , author=. Proceedings of the International Symposium on Human Factors and Ergonomics in Health Care , volume=. 2019 , organization=

  217. [256]

    Achiam, Josh and Adler, Steven and Agarwal, Sandhini and Ahmad, Lama and Akkaya, Ilge and Aleman, Florencia Leoni and Almeida, Diogo and Altenschmidt, Janko and Altman, Sam and Anadkat, Shyamal and others , journal=

  218. [257]

    Capabilities of

    Nori, Harsha and King, Nicholas and McKinney, Scott Mayer and Carignan, Dean and Horvitz, Eric , journal=. Capabilities of

  219. [258]

    Jama , volume=

    Accuracy of a generative artificial intelligence model in a complex diagnostic challenge , author=. Jama , volume=. 2023 , publisher=

  220. [259]

    Eriksen, Alexander V and M. Use of. NEJM AI , volume=. 2023 , publisher=

  221. [260]

    Capabilities of

    Antaki, Fares and Milad, Daniel and Chia, Mark A and Gigu. Capabilities of. British Journal of Ophthalmology , year=

  222. [261]

    Machine Learning for Health (ML4H) , pages=

    Med-flamingo: a multimodal medical few-shot learner , author=. Machine Learning for Health (ML4H) , pages=. 2023 , organization=

  223. [262]

    arXiv preprint arXiv:2305.09617 , year=

    Towards expert-level medical question answering with large language models , author=. arXiv preprint arXiv:2305.09617 , year=

  224. [263]

    arXiv preprint arXiv:2307.15343 , year=

    Med-halt: Medical domain hallucination test for large language models , author=. arXiv preprint arXiv:2307.15343 , year=

  225. [264]

    NPJ Digital Medicine , volume=

    Large language models propagate race-based medicine , author=. NPJ Digital Medicine , volume=. 2023 , publisher=

  226. [265]

    Advances in neural information processing systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=

  227. [266]

    NEJM AI , volume=

    Almanac—retrieval-augmented language models for clinical medicine , author=. NEJM AI , volume=. 2024 , publisher=

  228. [267]

    arXiv preprint arXiv:1701.06538 , year=

    Outrageously large neural networks: The sparsely-gated mixture-of-experts layer , author=. arXiv preprint arXiv:1701.06538 , year=

  229. [268]

    Journal of Machine Learning Research , volume=

    Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity , author=. Journal of Machine Learning Research , volume=

  230. [269]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=

    Randaugment: Practical automated data augmentation with a reduced search space , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops , pages=

  231. [270]

    arXiv preprint arXiv:2305.12031 , year=

    Clinical camel: An open-source expert-level medical language model with dialogue-based knowledge encoding , author=. arXiv preprint arXiv:2305.12031 , year=

  232. [271]

    2020 , publisher=

    Wagner, Patrick and Strodthoff, Nils and Bousseljot, Ralf-Dieter and Kreiseler, Dieter and Lunze, Fatima I and Samek, Wojciech and Schaeffter, Tobias , journal=. 2020 , publisher=

  233. [272]

    2024 , note =

    Image Challenge , author =. 2024 , note =

  234. [273]

    arXiv preprint arXiv:2307.03987 , year=

    A stitch in time saves nine: Detecting and mitigating hallucinations of llms by validating low-confidence generation , author=. arXiv preprint arXiv:2307.03987 , year=

  235. [274]

    Journal of the Medical Library Association , volume=

    When less is more: a practical approach to searching for evidence-based answers , author=. Journal of the Medical Library Association , volume=. 2002 , publisher=

  236. [275]

    The Bell system technical journal , volume=

    A mathematical theory of communication , author=. The Bell system technical journal , volume=. 1948 , publisher=

  237. [276]

    Nature Reviews Genetics , volume=

    Mining electronic health records: towards better research applications and clinical care , author=. Nature Reviews Genetics , volume=. 2012 , publisher=

  238. [277]

    Journal of the American Medical Informatics Association , volume=

    Extracting information from the text of electronic medical records to improve case detection: a systematic review , author=. Journal of the American Medical Informatics Association , volume=. 2016 , publisher=

  239. [278]

    arXiv preprint arXiv:2402.02008 , year=

    How well do LLMs cite relevant medical references? An evaluation framework and analyses , author=. arXiv preprint arXiv:2402.02008 , year=

  240. [279]

    arXiv preprint arXiv:2308.14089 , year=

    Medalign: A clinician-generated dataset for instruction following with electronic medical records , author=. arXiv preprint arXiv:2308.14089 , year=

  241. [280]

    Nature Biomedical Engineering , volume=

    Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging , author=. Nature Biomedical Engineering , volume=. 2023 , publisher=

  242. [281]

    NPJ digital medicine , volume=

    Considerations for addressing bias in artificial intelligence for health equity , author=. NPJ digital medicine , volume=. 2023 , publisher=

  243. [282]

    arXiv preprint arXiv:2403.12025 , year=

    A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models , author=. arXiv preprint arXiv:2403.12025 , year=

  244. [283]

    arXiv preprint arXiv:2308.01317 , year=

    ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders , author=. arXiv preprint arXiv:2308.01317 , year=

  245. [284]

    Machine Learning for Health , pages=

    Improving radiology report generation systems by removing hallucinated references to non-existent priors , author=. Machine Learning for Health , pages=. 2022 , organization=

  246. [285]

    2022 , publisher=

    Rajpurkar, Pranav and Chen, Emma and Banerjee, Oishi and Topol, Eric J , journal=. 2022 , publisher=

  247. [286]

    Journal of medical Internet research , volume=

    Information overload in emergency medicine physicians: a multisite case study exploring the causes, impact, and solutions in Four North England national health service trusts , author=. Journal of medical Internet research , volume=. 2020 , publisher=

  248. [287]

    arXiv e-prints , pages=

    Training Small Multimodal Models to Bridge Biomedical Competency Gap: A Case Study in Radiology Imaging , author=. arXiv e-prints , pages=

  249. [288]

    Journal of the American College of Surgeons , volume=

    Rationale and use of the critical view of safety in laparoscopic cholecystectomy , author=. Journal of the American College of Surgeons , volume=. 2010 , publisher=

  250. [289]

    Annals of surgery , volume=

    Causes and prevention of laparoscopic bile duct injuries: analysis of 252 cases from a human factors and cognitive psychology perspective , author=. Annals of surgery , volume=. 2003 , publisher=

  251. [290]

    IEEE transactions on medical imaging , volume=

    Endonet: a deep architecture for recognition tasks on laparoscopic videos , author=. IEEE transactions on medical imaging , volume=. 2016 , publisher=

  252. [291]

    Scientific Data , volume=

    Cholec80-CVS: An open dataset with an evaluation of Strasberg’s critical view of safety for AI , author=. Scientific Data , volume=. 2023 , publisher=

  253. [292]

    arXiv preprint arXiv:2106.10916 , year=

    Surgical data science for safe cholecystectomy: a protocol for segmentation of hepatocystic anatomy and assessment of the critical view of safety , author=. arXiv preprint arXiv:2106.10916 , year=

  254. [293]

    International journal of computer assisted radiology and surgery , volume=

    Weakly supervised convolutional LSTM approach for tool tracking in laparoscopic videos , author=. International journal of computer assisted radiology and surgery , volume=. 2019 , publisher=

  255. [294]

    2022 26th International Conference on Pattern Recognition (ICPR) , pages=

    Pixel-accurate segmentation of surgical tools based on bounding box annotations , author=. 2022 26th International Conference on Pattern Recognition (ICPR) , pages=. 2022 , organization=

  256. [295]

    Surgical Endoscopy , volume=

    Artificial intelligence for phase recognition in complex laparoscopic cholecystectomy , author=. Surgical Endoscopy , volume=. 2022 , publisher=

  257. [296]

    Endo3d: Online workflow analysis for endoscopic surgeries based on 3d cnn and lstm , author=. OR 2.0 Context-Aware Operating Theaters, Computer Assisted Robotic Endoscopy, Clinical Image-Based Procedures, and Skin Image Analysis: First International Workshop, OR 2.0 2018, 5th ...

  258. [297]

    Journal of the American College of Surgeons , volume=

    A simple effective method for generation of a permanent record of the critical view of safety during laparoscopic cholecystectomy by intraoperative “doublet” photography , author=. Journal of the American College of Surgeons , volume=. 2014 , publisher=

  259. [298]

    arXiv preprint arXiv:2403.10131 , year=

    RAFT: Adapting Language Model to Domain Specific RAG , author=. arXiv preprint arXiv:2403.10131 , year=

  260. [299]

    Yue, Xiang and Ni, Yuansheng and Zhang, Kai and Zheng, Tianyu and Liu, Ruoqi and Zhang, Ge and Stevens, Samuel and Jiang, Dongfu and Ren, Weiming and Sun, Yuxuan and others , journal=

  261. [300]

    ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

    Visual answer localization with cross-modal mutual knowledge transfer , author=. ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=. 2023 , organization=

Pith tools

Reviewed May 15, 2026 · model on record in the stance chip above.