A proprietary clinical RAG system, VITA, scores higher than GPT-5.4 and other frontier LLMs on English HealthBench questions, and roughly matches the newest GPT model when re-tested with a neutral judge.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
A proprietary clinical RAG system, VITA, scores higher than GPT-5.4 and other frontier LLMs on English HealthBench questions, and roughly matches the newest GPT model when re-tested with a neutral judge.