{"id":"a135a868-cad3-4038-827d-cffa504cae44","arxiv_id":"2509.19833","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SDG-POD, a 6,400-text benchmark with LLM-voted training labels and human test labels, shows current open LLMs reach only about 60-62 F1 on detecting whether SDG news signals progress or regression.","lead":"Researchers built a benchmark, SDG-POD, that tests whether a news excerpt about a UN Sustainable Development Goal describes progress, regression, or neither. Open language models score only around 60 percent F1, and fine-tuning on computer-generated examples gives modest but measurable gains.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test-set label noise is the load-bearing uncertainty: kappa as low as 0.38 and 480/576 single-annotator labels leave the 1.8-point headline gap and per-SDG rankings unsupported by the evidence as reported.","rationale":"I read the paper's central contribution as the SDG-POD benchmark plus the claim that synthetic data yields consistent gains. For either to stand, the human-annotated test set must be trustworthy. The reported inter-annotator agreement is moderate at best (kappa 0.38-0.69), and the protocol leaves most test items with a single annotator; this is the place where the argument is least secure. The reader's weakest_assumption identifies exactly this. I also considered the alternate concern that the 'synthetic data improves' claim is confounded with fine-tuning itself, since there is no non-synthetic fine-tuning control; that is real and should be fixed, but it is secondary because even a perfectly supported synthetic-data effect would not rescue the benchmark if the gold labels cannot support the reported gaps. The paper deserves credit for open code, honest disclosure of the low kappa and the Phi-4 epoch adjustment, and for framing the task as hard. The appropriate response is to keep the CONDITIONAL recommendation: the benchmark can be accepted conditionally on a reliability/robustness analysis of the test set. My read does not move the verdict, so verdict_should_be is UNCHANGED.","tokens_in":18121,"tokens_out":6336,"duration_ms":59943,"concrete_test":"Have two additional annotators independently re-annotate all 576 test items using the same guidelines, adjudicate disagreements with a third annotator, and recompute macro F1 for QWQ-32B FT and Phi-4 ZS on the adjudicated labels. If the original 1.8-point gap shrinks, disappears, or reverses, the reported benchmark comparisons are not robust to label noise. As a complementary analytical check, bootstrap the test set using label-flip probabilities derived from the pairwise disagreement rates in Tables 1-2; if the 95% interval for the QWQ-vs-Phi4 gap includes 0, the robustness claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SDG-POD is a valid benchmark and that fine-tuning on synthetic data yields robust gains, with QWQ-32B improving from 59.8 to 61.6 macro F1 and reducing positive/negative confusion. The load-bearing premise is that the human test labels are reliable enough to serve as ground truth. Section 4.3 reports pairwise Cohen's Kappa of 0.38-0.60 in Group 1 and 0.55-0.69 in Group 2 on only 48 shared texts per group. More importantly, 480 of the 576 test items are annotated by a single evaluator only; no reliability estimate is given for this majority of the test set. Every downstream number--the 1.8-point gap, the per-SDG rankings on 36 examples, the confusion-matrix claims in Tables 5-6, and the error-weighted F1s in Table 7--inherits this noise. The paper honestly discloses the low agreement, and McNemar was non-significant for neutral and negative classes, but the conclusions still treat the labels as exact. Without confidence intervals, bootstrap intervals, or noise-aware analysis, the observed differences are within the range that annotator disagreement alone could produce.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new task, SDG polarity detection (positive/neutral/negative with respect to a stated SDG), and presents SDG-POD, a 6,400-text benchmark built from OSDG excerpts. Training labels are produced by a five-LLM ensemble with majority voting and hand-written tie-break rules; the 576-item test set is annotated by six human annotators in two groups, with only 48 shared texts per group for agreement estimation. Six open-source models are evaluated in zero-shot and fine-tuned settings, and the paper claims that fine-tuning on the synthetic training data consistently improves performance, that QWQ-32B is the best model (macro F1 61.6), and that fine-tuning reduces critical positive/negative confusions, especially for SDG-9, SDG-12, and SDG-15.","tokens_in":18361,"tokens_out":4951,"duration_ms":40329,"significance":"If the central claims hold, the paper would provide a reusable evaluation harness and a low-cost LLM-voting annotation recipe for a resource-constrained, socially important domain. The authors are to be credited for releasing the codebase, for making the task definition explicit, and for honestly reporting the low inter-annotator agreement. However, the main empirical claims -- that synthetic training data improves robustness and that QWQ-32B is the best fine-tuned model -- are currently not established by the evidence as presented, because the experimental design confounds fine-tuning with label source, the comparison is not controlled, and the human-annotated test set is too noisy and too thinly annotated for the claimed differences to be reliable.","major_comments":[{"comment":"The test-set reliability is the load-bearing uncertainty. Cohen's Kappa values range from 0.38 to 0.60 in Group 1 and 0.55 to 0.69 in Group 2, yet these estimates come from only 48 shared texts per group. Moreover, 480 of the 576 test items are annotated by a single evaluator, with no reliability estimate for that majority. Every reported metric -- the 1.8-point gap between the best fine-tuned and best zero-shot model, the per-SDG rankings, and the confusion-matrix claims in Tables 5-6 -- inherits this label noise. The paper needs bootstrap confidence intervals, noise-aware evaluation, or a reliability analysis that treats the gold labels as uncertain; otherwise the stated conclusions are unsupported.","section":"§4.3, Tables 1-2"},{"comment":"The claim that synthetic data improves performance is confounded. The paper's hypothesis is that if fine-tuned models outperform their non-fine-tuned counterparts, this would indicate that the synthetic training data is valuable. But this comparison differs in two variables at once: fine-tuning itself and the provenance of the training labels. There is no control fine-tuned on non-synthetic (e.g., human-annotated) labels, nor a control using single-LLM labels, so the observed gains cannot be attributed to the synthetic label-generation method. Similarly, the claim that fine-tuning reduces critical positive/negative errors is based on comparing Phi-4 zero-shot (Table 5) with QWQ-32B fine-tuned (Table 6), which confounds model identity with training setting.","section":"§5, §6"},{"comment":"Phi-4 was fine-tuned for 10 epochs whereas all other models were fine-tuned for 5 epochs, explicitly 'in order to improve its classification metrics.' This makes cross-model comparisons of fine-tuned F1 scores unreliable and directly affects the reported 1.8-point gap between Phi-4 and QWQ-32B. If the goal is to compare model families under a standard protocol, the same number of epochs must be used or the results must be shown to be robust to epoch choice. The current presentation violates the standardised fine-tuning procedure stated at the start of §5.","section":"§6, Tables 4 and 7"},{"comment":"The per-SDG analysis is based on only 36 test examples per SDG, and the authors draw strong conclusions about SDG-specific rankings (e.g., SDG-9, SDG-12, SDG-15) without any uncertainty quantification. With a 3-class setup and 36 examples, a difference of a few instances can move a per-SDG F1 score by several points. Bootstrap intervals or error bars are necessary before per-SDG strengths can be claimed. The confusion matrices also need clearer formatting and row/column sums; as printed, several cell values do not obviously sum to the stated class totals.","section":"§6, Figure 1 and Tables 5-6"},{"comment":"The training labels themselves are generated by the authors' own five-LLM ensemble with hand-written Bronze tie-break rules, and there is no validation of these synthetic labels against human judgments. The paper presents no evidence that the ensemble's labels are accurate enough to serve as a training signal, beyond the aggregate downstream comparison that is already confounded. At minimum, a sample-based human evaluation of the synthetic labels, or a comparison with training on human labels of the same texts, is needed to support the claim that the proposed annotation recipe produces a 'high-quality resource for fine-tuning.'","section":"§4.2, §6"}],"minor_comments":[{"comment":"The title of the full text reads 'Polarity Detection of Sustainable Detection Goals,' which should be 'Sustainable Development Goals.' The abstract and title also say 'News Text,' but Section 4.1 describes the source as OSDG excerpts from reports, policy documents, and publication abstracts; please clarify the scope.","section":"Title and Abstract"},{"comment":"The caption refers to 'Phi4-8B,' but the model evaluated in the paper is PHI4-4B. Please correct the inconsistency.","section":"Figure 1 caption"},{"comment":"The sentence 'with a measured F1-score of 61.6%, an increase of 3.8 percentage points compared to the performance of the non-fine-tuned version of the same model' is ambiguous and appears to attribute the 61.6% figure to Phi-4 when it is actually QWQ-32B's score. Please rewrite to state clearly which model each number refers to.","section":"§6, paragraph after Table 4"},{"comment":"There is a typo: 'named olarity detection' should be 'named polarity detection.' Also, the related-work section would benefit from a clearer distinction between the proposed SDG-polarity task and the existing 'polarization analysis' literature, since the terms are used almost interchangeably.","section":"§2"},{"comment":"The description of the test-set annotation says a majority vote was used for the 48 shared texts, but it is not stated how the final label was assigned in the 2 'Bronze' cases (three different labels). The text mentions that neutral was chosen, but this appears only after the distribution of shared labels and should be stated more prominently.","section":"§4.3"},{"comment":"The confusion matrices are difficult to parse because the column headers are repeated across three blocks and the cell values are not visually aligned. Please reformat as standard confusion matrices and verify all row sums, as some cells appear inconsistent with the stated class sizes.","section":"Tables 5-6"}],"recommendation":"major_revision","confidential_remarks":"The paper has a potentially useful benchmark and annotation recipe, but the experimental protocol needs substantive tightening before the central claims can be accepted. The label-noise issue is particularly serious because the head-to-head differences are small and the gold labels are only weakly reliable. I would encourage the editor to ask for a revision that adds non-synthetic fine-tuning controls, consistent training epochs, and uncertainty-aware evaluation, even if that requires additional experiments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a genuinely new task and a useful benchmark, and the authors are transparent about several of its limits. But the two load-bearing claims—that the gold labels are solid enough for a benchmark, and that synthetic data 'consistently improves' performance—are not actually supported by the evidence as presented. It still deserves a serious referee.\n\nWhat's new: the three-way polarity label conditioned on a specific SDG is a meaningful gap in the literature, distinct from both plain SDG classification and aspect-based sentiment over national reports. The paper ships a 6,400-text dataset, a five-LLM majority-vote labeling procedure, open code, and an evaluation on open models only. The authors also report McNemar tests, disclose the Phi-4 epoch adjustment, and flag that the neutral and negative class improvements are not statistically significant. That is honest, and it earns real credit.\n\nThe soft spot that matters most is the test set. Pairwise Cohen's Kappa within the human group runs as low as 0.38, and 480 of 576 test items were annotated by a single person, so there is no reliability estimate for most of the gold labels. Every headline number—the 1.8-point macro-F1 gap between fine-tuned QWQ and zero-shot Phi-4, the per-SDG rankings on 36 examples, the confusion-matrix comparisons—inherits this noise, and the paper treats the labels as exact. A simple bootstrap or a noise-injected analysis would show how much of the reported difference remains; without that, the 'consistently improves' claim rests on a shaky base. The weighted F1 results with a tenfold penalty on positive/negative confusion show a much larger gap, which is more robust, but the cost weights are hand-chosen and the key confound is not addressed: fine-tuning on the synthetic labels is compared only to zero-shot, never to fine-tuning on a non-synthetic or human-labeled set. So you cannot separate the benefit of the synthetic data from the benefit of fine-tuning itself. Minor but real: the abstract says six models, the paper describes five; and Phi-4 was run for 10 epochs after seeing the test results.\n\nNone of these are fatal on their own. The task is worth having, the benchmark is a reasonable starting point, and the transparency suggests the authors will respond well to referee requests for a noise analysis and a control experiment. I would cite this for the task and dataset, not for the synthetic-data result. Send it to review, but with the explicit expectation that the central quantitative claims need to be made honest before publication.","headline":"SDG-POD is a real contribution, but the synthetic-data claim is confounded and the test-set noise makes the headline gaps unsupported as reported.","tokens_in":18963,"tokens_out":3612,"would_cite":true,"duration_ms":25803,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes SDG polarity detection—judging whether a news text signals progress, regression, or neutrality toward a specific Sustainable Development Goal—and introduces SDG-POD, a benchmark for it.","keywords":["SDG polarity detection","SDG-POD benchmark","large language models","synthetic data augmentation","multi-LLM majority voting","zero-shot vs fine-tuning","sustainability monitoring","text classification"],"falsifier":"Take the SDG-POD test set, have a fresh panel of annotators label the same 576 texts under the same guidelines, and re-run the comparison. If pairwise agreement on the shared items stays below 0.6 and the fine-tuned QWQ-32B no longer beats zero-shot Phi-4, the paper's central claim of consistent fine-tuning gains is not stable.","tokens_in":17926,"feed_emoji":"🌍","tokens_out":6116,"duration_ms":48332,"temperature":0.7,"pith_summary":"This paper introduces SDG-POD, a benchmark for a new task: given a news text and one of the UN's 17 Sustainable Development Goals, decide whether the text reports progress toward that goal, regression away from it, or neither. The central claim is that the task is genuinely hard for current open-source large language models, but that fine-tuning on a training set labelled by majority vote of five LLMs consistently improves performance. The best model, QWQ-32B, reaches a macro F1 of 61.6, up from 57.8 in zero-shot, and importantly reduces the severe error of swapping positive and negative labels. The paper also shows that synthetic training data can substitute for scarce human annotations in this domain. If the benchmark holds up, it gives sustainability researchers a reusable testbed and a low-cost recipe for building polarity detectors.","feed_headline":"Fine-tuned QWQ-32B tops SDG polarity test at 61.6 F1","feed_subtitle":"Synthetic training data from five LLMs lifts accuracy and cuts positive/negative confusion in sustainability monitoring.","key_machinery":"SDG-POD, a benchmark of 6,400 texts: 5,824 automatically labelled by a majority vote of five different LLMs and 576 independently labelled by human annotators. The load-bearing mechanism is the multi-LLM majority-voting pipeline that converts noisy individual predictions into silver training labels, combined with a cost-sensitive F1 evaluation that penalises positive/negative confusion.","core_discovery":"The paper's central discovery is that SDG polarity—whether a text reports progress toward, regression from, or neutrality about a given goal—can be learned from synthetic training data produced by majority-voting independent LLM annotations, improving on zero-shot prompting. Fine-tuned QWQ-32B reaches macro F1 61.6 versus 57.8 zero-shot; the best zero-shot model, Phi-4, scores 59.8. The gain is concentrated where it matters: fine-tuning reduces positive/negative label confusion, and under error-weighted F1 the gap widens from 1–2 points to 10–11 points. Per-SDG gains are largest on SDG-9, SDG-12, and SDG-15. The task remains challenging, making SDG-POD a demanding benchmark.","pith_inferences":["Editorial extension: because the test set is English and drawn from formal reports and policy documents, the polarity signal may not transfer to informal or multilingual text; a multilingual extension would reveal whether the task's difficulty is linguistic or conceptual.","Editorial extension: the reported 1.8-point gap between fine-tuned QWQ-32B and zero-shot Phi-4 sits close to the inter-annotator agreement range (Kappa 0.38–0.69), so a confidence-interval or noise-aware analysis is the natural next step for deciding whether the gap is stable.","Editorial extension: the same multi-LLM vote-and-fine-tune recipe could be applied to other low-resource policy-monitoring tasks, such as detecting whether corporate climate pledges represent progress or greenwashing.","Editorial extension: because synthetic-data fine-tuning reduces positive/negative confusion, a testable next step is to compare against a same-size human-labelled training set; similar or better F1 would isolate the value of annotation quality versus scale."],"forward_implications":["Fine-tuning on synthetic multi-LLM data beats zero-shot prompting: QWQ-32B moves from 57.8 to 61.6 macro F1, and Phi-4 from 59.8 to 61.3.","Fine-tuning reduces the most consequential errors: positive/negative label swaps are largely replaced by neutral misclassifications, and the error-weighted F1 gap widens to 10–11 points at high penalties.","SDG polarity is separable from sentiment: a celebratory tone can describe regression, and a bitter tone can describe progress, so sentiment classifiers cannot be reused directly.","The SDG-POD benchmark is reusable and open, providing a testbed for future models on a task where human annotations are scarce.","Zero-shot performance varies only modestly across models (57.7–59.8 F1), so the task is not yet saturated and synthetic fine-tuning is a practical route for deployment."],"fun_headline_variants":["Fine-tuned QWQ-32B with synthetic data tops SDG polarity test","Synthetic training data boosts LLM SDG progress-vs-regression detection","SDG-POD benchmark: fine-tuned LLMs cut positive/negative confusion","Zero-shot LLMs lag on SDG news polarity; fine-tuning closes gap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The 576 human-labelled test items are treated as ground truth; if those labels are unreliable—and the paper reports Cohen's Kappa as low as 0.38 among its annotators—then every model ranking and F1 gap inherits that unreliability.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned QWQ-32B with synthetic data tops SDG polarity test","Synthetic training data boosts LLM SDG progress-vs-regression detection","SDG-POD benchmark: fine-tuned LLMs cut positive/negative confusion","Zero-shot LLMs lag on SDG news polarity; fine-tuning closes gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1203,"prompt_tokens":769,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":348}},"tokens_in":513,"tokens_out":434,"duration_ms":18842,"temperature":1.0,"reasoning_tokens":348,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T15:19:40.394983+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the SDG-POD test set, have a fresh panel of annotators label the same 576 texts under the same guidelines, and re-run the comparison. If pairwise agreement on the shared items stays below 0.6 and the fine-tuned QWQ-32B no longer beats zero-shot Phi-4, the paper's central claim of consistent fine-tuning gains is not stable.","supporting_citations":[],"review_version":1}