{"id":"30e4b066-389c-44b1-871e-2b2890eb1730","arxiv_id":"1908.02322","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A Dutch news dataset with publisher-level and article-level partisanship labels, released on GitHub, despite low inter-annotator agreement.","lead":"A new Dutch news dataset labels over 100,000 articles by publisher partisanship and 766 articles by reader assessment. The labels show low agreement between annotators, so the resource's reliability is limited.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Publisher-level partisanship labels are validated only by a Spearman correlation of 0.21 (§3), so the 100K-article subset may carry little article-level partisanship signal.","rationale":"The reader and I converge on the validation of publisher-level labels. The abstract's headline resource is the 100K+ publisher-level labeled articles; the 766 article-level articles are a secondary component. If the publisher-level labels do not track article content, the central contribution as a partisanship benchmark depends on an assumption the paper itself tests and finds only weakly supported. The paper is transparent about the 0.21 correlation and lists limitations, and I credit it for releasing raw survey data, but transparency about a weak validation does not make the labels informative. The low inter-annotator alpha (0.18) is a separate concern; it further weakens the article-level subset, but the publisher-level validation is the load-bearing issue because it affects the large majority of the dataset. A simple AUC check with publisher labels as the only predictor would settle whether the large subset has any article-level partisanship signal. Because this concern is essentially the reader's weakest assumption and the verdict was already CONDITIONAL, I recommend no change to the verdict.","tokens_in":6326,"tokens_out":3654,"duration_ms":38521,"concrete_test":"Using the released raw survey data, derive article-level binary labels with an independent procedure that avoids the paper's post-hoc filtering (e.g., leave-one-annotator-out majority vote on all articles with at least two annotations). Then fit a logistic regression using only the publisher-level binary partisanship label as the sole feature to predict these article-level labels, and report AUC. If the AUC is close to 0.5 and the publisher/percentage Spearman correlation remains below roughly 0.3, the 100K publisher-level labels are not a usable partisanship signal for individual articles.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that more than 100K articles are labeled with partisanship via publisher-level labels. The paper's own validation of this paradigm in Section 3 is weak: it computes the Spearman correlation between audience-derived publisher partisanship scores (Table 2) and the percentage of article-level partisan labels per publisher (Table 7), obtaining 0.21. With only 11 publishers, this correlation is far from statistically significant. Moreover, the absolute scores all lie in a narrow band (−0.64 to 0.07), and article-level partisan rates are nearly flat across publishers (e.g., partisan VK is 27.1% while non-partisan AD is 24.1%, and regional non-partisan publishers range from 17.6% to 41.2%). Thus the binary publisher labels used for the 103,812 publisher-level articles may encode publisher identity rather than article partisanship. This substantially weakens the utility of the large half of the dataset as a partisanship benchmark. The article-level labels are not a reliable substitute either: Krippendorff's alpha is 0.18 even after post-hoc filtering, and the filtering steps (removing 'uninterested' and 'unreliable' annotators, discarding articles with fewer than three annotations, majority voting) are not independently validated. The strongest claim requires the publisher-level labels to carry signal about article content; the reported correlation gives little evidence that they do. The released raw survey data are a mitigating factor, but the advertised 'more than 100K publisher-level labeled articles' is the load-bearing part of the resource.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DpgMedia2019, a Dutch news dataset with two components: a publisher-level part of 103,812 articles from 11 DPG Media publishers, labeled partisan/non-partisan on the basis of the average self-reported political leaning of subscribers (Table 2), and an article-level part of 766 (elsewhere 776) articles labeled by subscribers through an internal survey that asked about perceived bias intensity, polarity, and pro/anti entities (Section 2.2, Appendix). The paper documents the collection pipeline, post-hoc filtering steps intended to improve annotator reliability, the final Krippendorff's alpha of 0.18, and a Spearman correlation of 0.21 between publisher-level partisanship scores and per-publisher article-level partisan percentages (Section 3). The authors also discuss limitations and suggest applications such as training a partisan news detector, and they state that raw survey data are released.","tokens_in":6597,"tokens_out":3497,"duration_ms":39690,"significance":"If the labels were valid, this would be the first Dutch-language benchmark for partisan news detection, and the combination of one hundred thousand publisher-level articles with a smaller article-level set would be a useful resource. The paper is transparent about its process and releases raw survey responses, which is a genuine strength: future work can re-derive labels with better methods. However, the advertised contribution is a dataset with partisanship labels, and the paper's own validation evidence is weak: the publisher-level label paradigm is supported only by a statistically insignificant Spearman correlation over 11 publishers, and the article-level labels rest on an inter-rater agreement of alpha = 0.18 after multiple post-hoc filtering steps. The resource is valuable as a collection with raw annotations, but the central claim of a reliable partisanship-labeled dataset is not yet established.","major_comments":[{"comment":"The publisher-level labeling paradigm is validated only by a Spearman correlation of 0.21 computed over 11 publishers between the absolute audience-based partisanship scores and the per-publisher percentage of article-level partisan labels. With n = 11, this correlation is not statistically significant, and the raw numbers in Tables 2 and 7 show why: the publisher scores lie in a narrow band from -0.64 to 0.07, and the per-publisher partisan rates are nearly flat (VK 27.1% versus AD 24.1%; regional publishers from 0% to 41.2%). This does not support the assumption that partisan publishers publish more partisan articles in this dataset, so the 103,812 publisher-level labels may encode publisher identity rather than article content. The authors should either supply independent validation (for example, a random sample of publisher-labeled articles judged by experts, or a transfer experiment showing article-level signal) or explicitly reframe the large component as an article collection with publisher-leaning scores rather than a partisanship-labeled benchmark.","section":"Section 3, Tables 2 and 7"},{"comment":"The article-level labels rest on very low inter-rater agreement: Krippendorff's alpha rises only from 0.142 to 0.180 after four filtering steps, including discarding 'uninterested' and 'unreliable' annotators, discarding articles with fewer than three annotations, majority voting, and removing articles where half or more annotations were 'impossible to decide'. The reliability score in equation (1) and the threshold on it are chosen post hoc based on the alpha they produce, so the final agreement estimate is not an independent measure of label quality. With alpha = 0.18, majority-voted labels from three or more untested annotators are close to chance-level agreement, and the paper provides no confidence interval, no per-class agreement, and no validation against expert labels. The authors should provide such evidence, or clearly restrict the claim of article-level partisanship labeling to the raw annotations rather than to the derived 766 labels.","section":"Section 2.2.2, Table 1"},{"comment":"The binary partition of publishers into partisan and non-partisan is not given a principled criterion. The computed scores are all slightly left-leaning (from -0.64 to 0.07), and the cutoff that places Het Parool (-0.434) into the partisan set and de Gelderlander (-0.245) into the non-partisan set is introduced as a decision without stating a rule or a pre-registered threshold. Because this cutoff determines the labels for all 103,812 publisher-level articles, its arbitrariness is load-bearing for the dataset's central claim. The paper should either present the continuous publisher scores as the primary publisher-level signal, with the binary labels clearly marked as one possible discretization, or justify the cutoff with an external criterion.","section":"Section 2.1.1, Table 2"}],"minor_comments":[{"comment":"The number of article-level labeled articles is given as 776 in the abstract but 766 in Table 1 and Section 2.2.2; this inconsistency should be corrected.","section":"Abstract, Table 1, Section 2.2.2"},{"comment":"There are several typographical errors: 'Krippendorf' should be 'Krippendorff', 'Spearsman' should be 'Spearman', 'dateset' in the Table 1 caption should be 'dataset', 'drabantsdagblad' in Table 2 should be 'brabantsdagblad', and publisher names are not consistently capitalized (for example 'destem' vs 'de Stem').","section":"Throughout"},{"comment":"The article length statistics do not state the unit of measurement; the authors should specify whether the values are word counts or character counts, and they should note that the publisher-level and article-level distributions have different standard deviations (387.5 vs 275.1) even though the means are close.","section":"Table 8"},{"comment":"The survey question Q1 asks about 'bevooroordeeld' (biased) rather than 'partisan' specifically, while the dataset is described throughout as a partisanship dataset; the paper should clarify the intended relationship between perceived bias and partisanship, since the two concepts are not identical.","section":"Section 2.2 and Appendix"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection because the raw survey data and the candid documentation of the design decisions have real value for the community. However, the advertised labels, especially the 100K publisher-level component, are not currently supported by the evidence in the paper; the authors need to either substantially strengthen the validation or reframe the contribution to match what is actually delivered. If they choose the latter route, the paper could become a solid resource-description paper rather than a partisanship benchmark paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuine resource paper — the first Dutch partisan news dataset, released on GitHub, with 103K publisher-level articles and about 770 article-level annotations plus the raw survey data. Its best feature is honesty: the limitations section names the small publisher samples, the narrow spread of partisanship scores, and the loosely defined annotation task. The paper does not hide its problems.\n\nThe problems are real and load-bearing. Article-level agreement is Krippendorff's alpha 0.18 after several post-hoc filtering steps (dropping 'uninterested' and 'unreliable' annotators, discarding articles with fewer than three annotations, majority voting). The filtering thresholds were chosen by watching the alpha, so 0.18 is the best case, not an independent estimate. The publisher-level half is validated only by a Spearman correlation of 0.21 between audience-derived publisher scores and article-level partisan percentages, computed over 11 publishers. That is not statistically significant, and the article-level rates are nearly flat across publishers: partisan VK sits at 27.1% while 'non-partisan' AD is 24.1%, and the regional non-partisan publishers range from 17.6% to 41.2%. The paper's own Section 1 premise — that partisan publishers publish more partisan articles — gets a weak test in Section 3 and mostly fails. So the 100K publisher-level labels likely encode publisher identity more than article content.\n\nMinor items: the abstract says 776 articles while Table 1 says 766; and the annotator pool overlaps the publishers being rated, though the low correlation between the two label sets at least shows they are not feeding each other.\n\nWhere this leaves the resource: too unreliable to serve as a partisanship benchmark on its own, but the raw survey data is a real asset for researchers who want to re-derive labels with better methods, and the dataset is a useful object for Dutch media studies and for studying disagreement in political annotation.\n\nRecommendation: send it to review. A serious referee should ask for sensitivity analysis of the filtering thresholds, a narrower scoping of what the publisher-level labels can support, and the count fix. The claims need conditioning, but the resource deserves to exist.","headline":"A genuine and honest dataset release — first Dutch partisan news corpus — but alpha 0.18 and Spearman 0.21 mean the labels are weakly validated and the claims need scoping.","tokens_in":7162,"tokens_out":3218,"would_cite":true,"duration_ms":31780,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents DpgMedia2019, a Dutch news dataset with more than 100,000 publisher-level partisanship labels and 766 article-level labels, built to support automated partisan-news detection in Dutch.","keywords":["Dutch news dataset","partisanship detection","news media bias","publisher-level labels","article-level annotation","crowdsourced annotation","political ideology","natural language processing"],"falsifier":"Take the 766 article-level labeled articles and compare each majority label with its publisher-level partisan/non-partisan label; if agreement is no better than always predicting non-partisan (the 74% base rate), the publisher-level layer does not carry the signal the dataset claims.","tokens_in":6106,"feed_emoji":"📰","tokens_out":9315,"duration_ms":95965,"temperature":0.7,"pith_summary":"The paper aims to supply a missing resource for Dutch-language media research: a large, public collection of news articles with explicit partisanship labels. The dataset has two layers: 103,812 articles labeled through the partisanship of their publisher, and 766 articles labeled by readers who answered survey questions about each article. If the resource works as intended, natural-language systems can train and test Dutch partisan-news detectors without starting from scratch. The paper also documents how the labels were collected and where the labeling process is noisy.","feed_headline":"Dutch news dataset labels 100K articles for partisanship","feed_subtitle":"A 766-article hand-labeled subset offers a check on the publisher-level shortcut.","key_machinery":"The mechanism is a two-layer labeling pipeline. The first layer uses an audience-based partisanship score: each publisher's score is the average of its subscribers' self-reported positions on a five-point left-right scale, and a threshold turns those scores into binary partisan or non-partisan publisher labels that propagate to all articles from that publisher. The second layer is a subscriber survey in which readers rate each article's bias intensity, polarity, and target; a reliability score based on free-text answers, filters for too-few annotations, and majority voting convert the raw ratings into article-level partisan labels. The paper then compares the two layers with a Spearman correlation, reporting 0.21, as a check on the outer labeling assumption.","core_discovery":"The central claim is that DpgMedia2019 fills that gap. It contains 103,812 unique articles from 11 Dutch publishers, split almost evenly into 52,873 publisher-level partisan labels and 50,939 non-partisan labels, plus 766 crowd-annotated articles of which 201 are partisan. Publisher partisanship was derived by averaging the self-reported political leaning of readers per publisher; de Volkskrant, Trouw, and Het Parool were treated as partisan and the remaining eight as non-partisan. The article-level labels came from an internal subscriber survey, filtered by a reliability score, with majority voting across at least three annotations per article. The authors present the dataset as a foundation for a partisan-news detector, to be used, for example, by training on one label layer and testing on the other.","pith_inferences":["The near-chance gap in the publisher-article comparison suggests the publisher-level label is better treated as a weak prior than as a gold label; article-level evaluation may be necessary for credible detector claims.","With only 26% of article-level examples partisan, accuracy is a misleading metric; precision, recall, and calibration by annotator self-reported leaning would be more informative.","The regional publishers show a spread in partisan percentages, so the dataset could support studies of regional variation in Dutch news partisanship rather than only national titles.","The released raw ratings could be re-analyzed with models that treat each annotator's political leaning as a parameter, which might yield more consistent labels than the current majority vote."],"forward_implications":["A Dutch partisan-news detector can be trained on the 100K publisher-level labels and evaluated on the 766 article-level labels, giving an in-language benchmark.","The publisher-level portion can serve as unlabeled data in semi-supervised learning while the article-level portion provides supervision.","Because the raw survey responses are released, alternative label-aggregation schemes can be tested against the reported majority-vote labels.","The reported low correlation between label layers means any model trained on publisher labels should be validated on article labels before being trusted."],"supporting_citations":[{"why":"Precedent for labeling articles through publisher partisanship, the shortcut the first dataset layer adopts.","marker":"Potthast et al., 2018"},{"why":"Earlier use of publisher-level ideology labels for news articles, cited as support for the labeling paradigm.","marker":"Kulkarni et al., 2018"},{"why":"Provides the hyperpartisan-news annotation questions that the Dutch survey adapted into article-level labels.","marker":"Kiesel et al., 2019"},{"why":"External finding on Dutch media partisanship used to sanity-check the audience-derived publisher scores.","marker":"Center, 2018a"}],"fun_headline_variants":["Dutch news partisanship dataset: 100K articles, 766 hand-checked","New Dutch news corpus for partisanship detection with 100K labels","Partisan news labels for 11 Dutch publishers, plus 766 crowd-checked","DpgMedia2019: Dutch partisanship dataset with publisher and article labels","Dutch news partisanship: 100K publisher labels, 766 verified articles"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a publisher's subscriber-averaged political leaning tracks the partisanship of its individual articles closely enough to label more than 100K articles with more signal than noise; the paper's own check of this premise is a Spearman correlation of 0.21.","fun_headline_variants_meta":{"raw":{"variants":["Dutch news partisanship dataset: 100K articles, 766 hand-checked","New Dutch news corpus for partisanship detection with 100K labels","Partisan news labels for 11 Dutch publishers, plus 766 crowd-checked","DpgMedia2019: Dutch partisanship dataset with publisher and article labels","Dutch news partisanship: 100K publisher labels, 766 verified articles"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000181,"raw_usage":{"total_tokens":1208,"prompt_tokens":748,"completion_tokens":460,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":364,"completion_tokens_details":{"reasoning_tokens":355}},"tokens_in":364,"tokens_out":460,"duration_ms":4475,"temperature":1.0,"reasoning_tokens":355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:46:58.511186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the 766 article-level labeled articles and compare each majority label with its publisher-level partisan/non-partisan label; if agreement is no better than always predicting non-partisan (the 74% base rate), the publisher-level layer does not carry the signal the dataset claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent for labeling articles through publisher partisanship, the shortcut the first dataset layer adopts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier use of publisher-level ideology labels for news articles, cited as support for the labeling paradigm."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the hyperpartisan-news annotation questions that the Dutch survey adapted into article-level labels."}],"review_version":1}