Pith. sign in

REVIEW 20 cited by

MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.14371 v1 pith:7GCHZUZO submitted 2022-03-27 cs.CL cs.AIcs.LG

MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

classification cs.CL cs.AIcs.LG
keywords medicalquestionansweringdatasetentranceexamlarge-scalemedmcqa
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper introduces MedMCQA, a new large-scale, Multiple-Choice Question Answering (MCQA) dataset designed to address real-world medical entrance exam questions. More than 194k high-quality AIIMS \& NEET PG entrance exam MCQs covering 2.4k healthcare topics and 21 medical subjects are collected with an average token length of 12.77 and high topical diversity. Each sample contains a question, correct answer(s), and other options which requires a deeper language understanding as it tests the 10+ reasoning abilities of a model across a wide range of medical subjects \& topics. A detailed explanation of the solution, along with the above information, is provided in this study.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences?

    cs.AI 2026-04 unverdicted novelty 7.0

    LLMs predict outcomes of real scientific experiments at 14-26% accuracy, comparable to human experts, but lack calibration on prediction reliability while humans demonstrate strong calibration.

  2. Automatic Replication of LLM Mistakes in Medical Conversations

    cs.CL 2025-12 unverdicted novelty 7.0

    MedMistake automatically generates 3,390 single-shot QA pairs capturing LLM mistakes in medical conversations, with expert validation on a 211-question subset showing performance differences among 12 frontier models.

  3. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

    cs.AI 2026-07 conditional novelty 6.5

    A real multimodal Chinese online-consultation benchmark of 5,620 cases finds frontier LLMs below physicians, with safety-sensitive error avoidance as the main gap.

  4. Reference-Free Evaluation of Reasoning in Open-Ended Question Answering

    cs.CL 2026-07 conditional novelty 6.0

    An NLI-hypergraph audit with deterministic AND–OR search labels LLM reasoning segments as supported, unsupported, or orphaned, improving balanced F1 over LLM-as-judge on a new 40-case clinical benchmark.

  5. Cura 1T: Specialized Model for Agentic Healthcare

    cs.AI 2026-07 conditional novelty 6.0

    A healthcare LLM trained with an agent-driven loop that converts benchmark failures into new training data outperforms frontier baselines on most healthcare tests while staying competitive on general reasoning.

  6. MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

    cs.AI 2026-07 conditional novelty 6.0

    On a new benchmark of 5,620 real multimodal online consultations, top LLMs trail the original physicians mainly because they trigger more unsafe or unsupported negative criteria.

  7. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

    cs.CL 2026-07 unverdicted novelty 6.0

    LLM evaluators reach clinician-level agreement on a new German medical benchmark but fail to abstain on difficult items and show lineage-dependent scoring biases.

  8. Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

    cs.CL 2026-07 conditional novelty 6.0

    A new German open-response medical benchmark shows LLM judges can match physician agreement (κ≈0.69 vs 0.71) yet lack clinical caution and show lineage-dependent scoring bias.

  9. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    cs.AI 2026-06 conditional novelty 6.0

    Systematic tests of LLM contamination detectors across 27 models show frequent failures from distribution shift and scale, concluding statistical methods cannot replace transparent data provenance.

  10. The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

    cs.AI 2026-06 conditional novelty 6.0

    Empirical evaluation across 25 LLMs shows contamination detection methods achieve correct outcomes in only 201 of 335 cases, exposing failure modes from distribution shift and benchmark scale.

  11. CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

    cs.CL 2026-05 unverdicted novelty 6.0

    CHI-Bench shows current AI agents achieve at most 28% success on long-horizon healthcare workflows that require dense policy adherence, multi-role handoffs, and multi-turn interactions.

  12. Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation

    cs.AI 2026-05 unverdicted novelty 6.0

    Self-evolving LLM agents exhibit capability erosion under continual adaptation, which Capability-Preserving Evolution mitigates by raising retained simple-task performance from 41.8% to 52.8% in workflow evolution und...

  13. Compared to What? Baselines and Metrics for Counterfactual Prompting

    cs.CL 2026-05 conditional novelty 6.0

    Counterfactual prompting effects on LLMs are often indistinguishable from those caused by meaning-preserving paraphrases, causing most previously reported demographic sensitivities to disappear under proper statistica...

  14. Eliciting Medical Reasoning with Knowledge-enhanced Data Synthesis: A Semi-Supervised Reinforcement Learning Approach

    cs.LG 2026-04 unverdicted novelty 6.0

    MedSSR improves LLM medical reasoning on rare diseases by up to 5.93% through knowledge-enhanced question synthesis and semi-supervised RL with self-generated pseudo-labels.

  15. Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA

    cs.CL 2026-06 unverdicted novelty 5.0

    Empirical comparison across model families finds SFT often sufficient for MCQA while CPT aids OEQA metrics in French medical LLM adaptation, with cross-lingual transfer observed.

  16. Generalistic or Specific Embeddings, Which is Better? An Empirical Study on Search for Clinical Coding in Non-English Languages

    cs.CL 2026-05 unverdicted novelty 5.0

    Fine-tuning a Spanish biomedical encoder on Gemini-generated synthetic data for multiple languages yields a bi-encoder that matches or exceeds BioBERT-ST on clinical code retrieval metrics, with further gains from cro...

  17. Active Evidence-Seeking and Diagnostic Reasoning in Large Language Models for Clinical Decision Support

    cs.AI 2026-05 unverdicted novelty 5.0

    Multi-turn evidence seeking reduces LLM diagnostic accuracy by 12.75% and supporting-evidence quality by 24.36% versus full-context evaluation in a new OSCE-inspired benchmark across 468 cases and 15 models.

  18. Claim-Selective Certification for High-Risk Medical Retrieval-Augmented Generation

    cs.CL 2026-05 unverdicted novelty 5.0

    Claim-selective certification decomposes medical RAG responses into verifiable claims scored against retrieved evidence and mapped via an intent-aware selector to actions, reporting zero UCCR and action accuracy of 0....

  19. Galactica: A Large Language Model for Science

    cs.CL 2022-11 unverdicted novelty 5.0

    Galactica, a science-specialized LLM, reports higher scores than GPT-3, Chinchilla, and PaLM on LaTeX knowledge, mathematical reasoning, and medical QA benchmarks while outperforming general models on BIG-bench.

  20. MDIA: A Multi-Agent Diagnostic Intelligence Pipeline on HealthBench Professional

    cs.AI 2026-05 unverdicted novelty 4.0

    MDIA, a specialty-routed 7-node multi-agent system, reports 0.6272 accuracy on 525 HealthBench Professional cases using GPT-5.4, outperforming the ChatGPT for Clinicians baseline by 3.72 points and attributing the lif...