VB-Score shows three major LLMs have severe failures in medical entity recognition and factual consistency, with 13.8% lower performance on chronic conditions affecting older and minority groups, indicating condition-based algorithmic discrimination.
Calvert, Alastair K
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Analysis of benchmark gaps in low-resource settings leads to a proposed shared reporting framework that combines task performance with deployment conditions and uses one-page benchmark cards.
citing papers explorer
-
Beyond Semantic Similarity: A Component-Wise Evaluation Framework for Medical Question Answering Systems with Health Equity Implications
VB-Score shows three major LLMs have severe failures in medical entity recognition and factual consistency, with 13.8% lower performance on chronic conditions affecting older and minority groups, indicating condition-based algorithmic discrimination.
-
Benchmarking AI for low-resource contexts: Thinking beyond leaderboards
Analysis of benchmark gaps in low-resource settings leads to a proposed shared reporting framework that combines task performance with deployment conditions and uses one-page benchmark cards.