A lightweight multilingual QA smoke-test suite with a 52-item English core and an LLM-based generator is shown to reflect model size and language performance differences in seconds.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
A lightweight multilingual QA smoke-test suite with a 52-item English core and an LLM-based generator is shown to reflect model size and language performance differences in seconds.