A formalization of benchmarkless LLM safety scoring validated via an instrumental-validity chain of contrast separation, target variance dominance, and rerun stability, demonstrated on Norwegian scenarios.
Proceedings of the 24th Nordic Conference on Computational Linguistics (NoDaLiDa) , year =
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
Three BERT models are further pre-trained on Norwegian clinical notes and discharge summaries, then shown to outperform their base models on synthetic clinical benchmarks and real-world tasks.
citing papers explorer
-
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
A formalization of benchmarkless LLM safety scoring validated via an instrumental-validity chain of contrast separation, target variance dominance, and rerun stability, demonstrated on Norwegian scenarios.
-
KliniskVestBERT: BERT Model Specialised to Norwegian Clinical Texts
Three BERT models are further pre-trained on Norwegian clinical notes and discharge summaries, then shown to outperform their base models on synthetic clinical benchmarks and real-world tasks.