{"id":"de14f356-bad5-4a86-8007-4ef9e2dbfe29","arxiv_id":"2501.11128","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Four Norwegian QA datasets covering both written standards are released, and 11 language models are benchmarked on them.","lead":"This paper presents four new Norwegian question-answering datasets, with over 10,000 question-answer pairs in both Bokmål and Nynorsk, covering general knowledge, commonsense, truthfulness, and Norway-specific trivia. Researchers can use these public datasets and the included evaluation of 11 language models to measure how well AI systems handle Norwegian in multiple task formats.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Localization may change task difficulty; without IAA or human baselines, the adapted datasets' construct validity is unverified.","rationale":"The paper is a useful resource contribution with transparent annotation procedures, public release, and a broad evaluation of 11 models. The reader's CONDITIONAL verdict already captures the main risk: the adaptation pipeline may change what the datasets measure, and the lack of inter-annotator agreement and human baselines leaves this risk unquantified. My stress-test agrees with that reading and adds two concrete pieces of evidence available in the paper itself: Section 3.6 reports that the Norwegian items became shorter and less complex, and Section 7 acknowledges single-annotator validation without IAA. These are not fatal flaws, because the datasets are still usable and the limitations are disclosed, but they do mean the strong empirical claims about model abilities should be treated as provisional. I therefore recommend keeping the verdict at CONDITIONAL rather than upgrading to ACCEPT.","tokens_in":13739,"tokens_out":6238,"duration_ms":61818,"concrete_test":"Sample 100 items from each of NorOpenBookQA, NorCommonSenseQA, and NorTruthfulQA-MC, stratified by the translation-versus-creative-adaptation flags available in the released dataset fields. Have two fresh native Norwegian speakers, not involved in the original annotation, answer these items following the original English dataset protocols, and compare their accuracy to published human baselines on the corresponding English datasets (e.g., OpenBookQA and CommonSenseQA human accuracy). Separately, for NorOpenBookQA, run the same LMs with and without the provided 'fact' in the prompt. If Norwegian human accuracy is substantially above the English human baselines, or if LM accuracy without the fact is close to accuracy with it, the localization changed item difficulty and the datasets do not isolate the intended construct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central value of this resource is that NorOpenBookQA, NorCommonSenseQA, and NorTruthfulQA measure world knowledge, commonsense reasoning, and truthfulness in Norwegian. That claim depends on the adaptation in Section 3.1.1 preserving what the English originals measure. Annotators were allowed to 'localize' and 'creatively adapt' items, not merely translate them. Section 3.6's own manual comparison reports that the Norwegian questions are 'somewhat shorter' and 'less complex' than the English ones, e.g., 'Hvor kommer kumelk fra?' (Where does cow milk come from?). This is internal evidence that adaptation can lower the cognitive demand of an item. Because Section 7 explicitly states that each example was validated by a single annotator, no inter-annotator agreement was computed, and no human baseline was collected, there is no quantitative check that the Norwegian datasets still require the same knowledge or reasoning as OpenBookQA, CommonSenseQA, and TruthfulQA. If the items have become easier or more surface-cue driven, the paper's empirical conclusions ('LMs struggle most with commonsense reasoning', 'often untruthful', 'better in NB than NN') would reflect dataset artifacts rather than the intended abilities. This is load-bearing because the datasets are proposed as benchmarks for precisely those abilities.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents four new Norwegian QA datasets: NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA (multiple-choice and generation), and NRK-Quiz-QA, covering both Bokmål and Nynorsk. The datasets are created by native-speaker annotators through manual translation and localization of OpenBookQA, CommonSenseQA, and TruthfulQA, together with newly written examples and adapted material from NRK quizzes. The authors describe the annotation pipeline, report dataset statistics, and evaluate 11 language models in zero- and few-shot regimes. The main empirical findings are that most models score higher in Bokmål than Nynorsk, that commonsense reasoning is the weakest skill, and that models often reproduce human falsehoods. All datasets and annotation materials are publicly released.","tokens_in":13954,"tokens_out":5435,"duration_ms":51144,"significance":"If the resource is valid, it makes a substantial contribution: it is the first Norwegian QA collection covering both written standards, with multiple task formats and Norwegian-specific content, and it provides a reusable evaluation harness. The paper is commendably transparent about its limitations, including the 80% curation rate, lack of inter-annotator agreement, and absence of human baselines. The public release under a permissive license and the integration with NorEval are concrete strengths that will benefit the community. However, the strength of the empirical conclusions is currently limited by validation gaps; the resource itself is valuable, but the evidence that the adapted datasets measure the intended constructs is incomplete.","major_comments":[{"comment":"The paper's central empirical claims—that LMs struggle most with commonsense reasoning and are often untruthful in Norwegian—presuppose that NorCommonSenseQA and NorTruthfulQA measure the same constructs as their English source datasets. Section 3.1.1 explicitly permits annotators to localize and creatively adapt items, and Section 3.6 reports that the Norwegian questions are 'somewhat shorter' and 'less complex' than the English originals (e.g., 'Hvor kommer kumelk fra?'). This is internal evidence that adaptation can lower the cognitive demand of individual items. Because Section 7 states that only 80% of examples were curated, each by a single annotator, with no inter-annotator agreement, and no human baseline was collected, there is currently no quantitative check on construct validity. I ask the authors to add (i) a human baseline on a random sample, (ii) an item-level difficulty comparison with the English originals (e.g., correlation of per-item model accuracy), and (iii) a report of how many items were translated, localized, or newly created, with an analysis of whether localization changes answer distributions.","section":"§3.1.1, §3.6, §7"},{"comment":"The Nynorsk subsets are very small: NorOpenBookQA has 253 NN examples, NorCommonSenseQA has 95, and NorTruthfulQA Multiple-choice has 57. Claims such as 'most LMs perform better in NB than NN' and the NN-specific observations in Section 5 (e.g., only three models surpass 40% on NorCommonSenseQA NN) rest on samples where a difference of 5–8 percentage points is within binomial noise. The authors should report confidence intervals or significance tests for the NB–NN comparisons, and should soften or qualify the corresponding conclusions, particularly for datasets with fewer than roughly 100 NN examples.","section":"Table 2, §5"},{"comment":"The evaluation protocol reports the maximum accuracy and ROUGE-L across the pool of 50 prompts for each model. Maximum-of-prompts is an optimistic summary and can make model rankings sensitive to a single prompt; it is also not accompanied by any variance estimate. Since the paper's empirical conclusions are based on these numbers, I request that the authors also report mean and standard deviation (or per-prompt distributions), and that they fix and document the random seed used for k-shot demonstration sampling.","section":"§4 Result Aggregation"}],"minor_comments":[{"comment":"The dataset name is spelled inconsistently as 'NorCommonsenseQA' and 'NorCommonSenseQA'; please unify to a single spelling.","section":"Throughout"},{"comment":"The name 'NortruthfulQA' appears with a lowercase 't' at the start of Section 3.1; use 'NorTruthfulQA' consistently throughout.","section":"§3.1"},{"comment":"The manual comparison of 100 examples should state which datasets and annotators were involved and how the sample was drawn, so readers can interpret the claim about question length and complexity.","section":"§3.6"},{"comment":"The k-shot demonstration examples are sampled randomly without a stated seed; this makes the few-shot results non-reproducible as reported.","section":"§4"},{"comment":"The column header for NorOpenBookQA appears to duplicate 'NB'; check that the two k-shot blocks are labeled unambiguously.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid resource contribution: the dataset release is genuine, the annotation process is described in detail, and the authors are transparent about limitations. My major_revision recommendation is driven by the gap between the strength of the empirical conclusions and the lack of validation (human baselines, inter-annotator agreement, and adequate Nynorsk sample sizes). These issues are addressable through additional experiments and more cautious wording, and I would be comfortable with acceptance after they are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper gives Norwegian NLP something it didn't have: a public QA suite covering both Bokmål and Nynorsk, with world knowledge, commonsense, truthfulness, and a Norway-specific quiz set. That's a real gap, and the authors have done the work — 10.5k pairs, native speakers, a documented two-stage pipeline, and public release. NRK-Quiz-QA in particular is a genuinely new resource, not a translation.\n\nThe strongest part is the resource itself and the careful description of how it was built. The adaptation method is standard, but the execution is concrete: guidelines, training phase, per-example curation status on HuggingFace. The comparison table makes the novelty clear.\n\nThe soft spots are where the stress-test note lands. Section 3.6 reports that the Norwegian questions are 'somewhat shorter' and 'less complex' than the English originals, with examples like 'Hvor kommer kumelk fra?' (Where does cow milk come from?). That's internal evidence that localization can lower the cognitive demand of an item. Since only 80% of examples got single-annotator curation, with no inter-annotator agreement and no human baseline, there is no quantitative check that the adapted datasets still measure the same constructs as OpenBookQA, CommonSenseQA, and TruthfulQA. That matters because the evaluation section draws conclusions like 'LMs struggle most with commonsense reasoning' and 'are often untruthful.' Those conclusions could reflect item difficulty shifts rather than the intended abilities. The paper itself acknowledges all of this in Section 7, which is honest, but it means the empirical rankings should be read as suggestive, not definitive.\n\nTwo more minor points: the Nynorsk subsets are very small (95 examples for NorCommonSenseQA, 57 for NorTruthfulQA MC), so NB-vs-NN comparisons are fragile. And reporting maximum accuracy across 50 prompts without error bars or any sense of variance makes model differences look firmer than they are.\n\nWho is this for? Anyone building or evaluating Norwegian LMs, and people who care about how to adapt English benchmarks responsibly. It deserves a serious referee. My recommendation: send it to peer review, and ask the authors to either add human baselines and agreement metrics or explicitly soften the empirical claims in the abstract and conclusion. The resource is worth having either way.","headline":"A solid, genuinely useful Norwegian QA resource that fills a real gap, but the empirical model-ranking claims should be read with the paper's own caveats about localization and annotation quality in mind.","tokens_in":14527,"tokens_out":2534,"would_cite":true,"duration_ms":22720,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new suite of four Norwegian QA datasets, built by native speakers, covers both Bokmål and Nynorsk and reveals that language models perform worse in Nynorsk, struggle with commonsense reasoning, and often repeat falsehoods.","keywords":["Norwegian question answering","Bokmål","Nynorsk","commonsense reasoning","truthfulness","language model evaluation","dataset creation","multiple-choice QA"],"falsifier":"Have a group of bilingual Norwegian–English speakers answer a random sample of the original English OpenBookQA and CommonSenseQA items and the corresponding Norwegian versions, and compare human accuracy and item difficulty; if accuracy or difficulty diverge sharply, the Norwegian datasets are not measuring the same constructs. Also, double-annotate a sample of the curated items to measure inter-annotator agreement on quality; low agreement would weaken the curation claim.","tokens_in":13558,"feed_emoji":"🇳🇴","tokens_out":5388,"duration_ms":48147,"temperature":0.7,"pith_summary":"This paper introduces four new Norwegian question-answering datasets — NorOpenBookQA, NorCommonSenseQA, NorTruthfulQA, and NRK-Quiz-QA — that together cover world knowledge, commonsense reasoning, truthfulness, and Norway-specific knowledge in both written standards of Norwegian, Bokmål and Nynorsk. The 10.5k question-answer pairs were created by a team of native speakers through manual translation, localization, and creative adaptation of English benchmarks, plus curation of NRK quiz data. Evaluating 11 language models, the authors find that most perform better in Bokmål than Nynorsk, struggle most on commonsense reasoning, and often generate or select false answers. If the datasets measure what they claim, Norwegian language modeling now has a benchmark that goes beyond reading comprehension and can track progress in reasoning and truthfulness for both written standards.","feed_headline":"Four Norwegian QA datasets put language models to the test","feed_subtitle":"Native speakers wrote 10.5k questions in Bokmål and Nynorsk; models lag on commonsense and truthfulness.","key_machinery":"The central object is the collection itself, built by a two-stage annotation process: first, native-speaker annotators manually translated, localized, and creatively rewrote items from OpenBookQA, CommonSenseQA, and TruthfulQA; second, a curation stage filtered low-quality examples and checked spelling and grammar. NRK-Quiz-QA was assembled from more than 500 quizzes from the Norwegian public broadcaster, with temporal references adjusted and image-dependent items removed. The evaluation uses multiple-choice scoring by token probability for three datasets and free-form generation scored by ROUGE-L for the generation subset, across a pool of prompts per language standard.","core_discovery":"On the paper's own terms: Norwegian QA resources have been limited to extractive or translated reading-comprehension datasets in Bokmål; this collection is the first to cover both official written standards and to target commonsense reasoning, truthfulness, and Norwegian-specific knowledge, with items written by native speakers rather than machine-translated. The reported evaluations show consistent performance gaps that suggest these abilities are not yet well supported in Norwegian language models.","pith_inferences":["Because localization can change question difficulty, the claim that models struggle with commonsense reasoning could be an artifact of changed item difficulty; a bilingual human-answer study comparing item-level difficulty against the English originals would test this directly.","The Nynorsk subsets are small (e.g., 95 examples for NorCommonSenseQA and 57 for NorTruthfulQA multiple-choice), so NN score differences should be read with wide error bars; expanding NN coverage is a natural next step.","The NRK quiz items are zeitgeist-dependent and may leak into pretraining corpora; a contamination check on those items would harden the benchmark.","The NorOpenBookQA training split enables future instruction tuning or continued pretraining on Norwegian world-knowledge data, which the authors do not explore."],"forward_implications":["Norwegian language models can now be benchmarked on commonsense and truthfulness separately from reading comprehension, in both Bokmål and Nynorsk.","The consistent Bokmål-over-Nynorsk gap identifies minority-language coverage as a specific weakness to target in pretraining and evaluation.","The low truthfulness scores indicate that Norwegian models reproduce human misconceptions, so mitigation efforts need Norwegian-specific data.","The few-shot results on NorOpenBookQA suggest that adding demonstrations is not monotonic; 4-shot often beats 16-shot, informing practical evaluation protocols.","The public release under a permissive license lets other researchers reuse the datasets for training and cross-lingual comparison."],"supporting_citations":[{"why":"Source of OpenBookQA, supplying the original elementary-level science questions and structure that NorOpenBookQA translates and localizes.","marker":"Mihaylov et al., 2018"},{"why":"Source of CommonSenseQA, providing the commonsense multiple-choice items that NorCommonSenseQA adapts.","marker":"Talmor et al., 2019"},{"why":"Source of TruthfulQA, providing the truthfulness questions and answer sets that NorTruthfulQA translates and localizes.","marker":"Lin et al., 2022"},{"why":"Existing Norwegian QA resource (Belebele) that this collection compares against and extends beyond.","marker":"Bandarkar et al., 2024"},{"why":"Existing Norwegian subset of MKQA, used as a comparison point for world-knowledge QA.","marker":"Longpre et al., 2021"},{"why":"NorQuAD, the existing Norwegian extractive QA dataset, used as a baseline for reading comprehension.","marker":"Ivanova et al., 2023"},{"why":"Existing Norwegian QA resources and language models from NLEBench+NorGLM, used for comparison and as evaluated models.","marker":"Liu et al., 2024"}],"fun_headline_variants":["Norwegian QA suite exposes model gaps in commonsense and truthfulness","First Norwegian QA datasets in both written standards, written by natives","Models lag on commonsense and truthfulness in new Norwegian QA suite","10k native-written Norwegian questions challenge LMs on truthfulness","New Norwegian QA sets: first to cover Nynorsk and Bokmål with native questions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline that translates, localizes, and creatively rewrites English items into Norwegian preserves the abilities the original datasets measure, and the single-annotator curation of 80% of examples is sufficient to ensure item quality.","fun_headline_variants_meta":{"raw":{"variants":["Norwegian QA suite exposes model gaps in commonsense and truthfulness","First Norwegian QA datasets in both written standards, written by natives","Models lag on commonsense and truthfulness in new Norwegian QA suite","10k native-written Norwegian questions challenge LMs on truthfulness","New Norwegian QA sets: first to cover Nynorsk and Bokmål with native questions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00068,"raw_usage":{"total_tokens":3003,"prompt_tokens":773,"completion_tokens":2230,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":389,"completion_tokens_details":{"reasoning_tokens":2138}},"tokens_in":389,"tokens_out":2230,"duration_ms":14272,"temperature":1.0,"reasoning_tokens":2138,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:36:15.000119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a group of bilingual Norwegian–English speakers answer a random sample of the original English OpenBookQA and CommonSenseQA items and the corresponding Norwegian versions, and compare human accuracy and item difficulty; if accuracy or difficulty diverge sharply, the Norwegian datasets are not measuring the same constructs. Also, double-annotate a sample of the curated items to measure inter-annotator agreement on quality; low agreement would weaken the curation claim.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Existing Norwegian QA resource (Belebele) that this collection compares against and extends beyond."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"NorQuAD, the existing Norwegian extractive QA dataset, used as a baseline for reading comprehension."}],"review_version":1}