{"id":"fa0987a4-e1a3-4096-9a87-402659c6fb6a","arxiv_id":"2506.02573","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"IndoSafety, a culturally grounded safety benchmark for five Indonesian language varieties, shows unsafe response rates up to 40% in regional models and demonstrates that safety tuning on formal Indonesian transfers to local languages.","lead":"The authors created IndoSafety, a human-verified safety evaluation dataset of 2,500 prompts spanning formal and colloquial Indonesian plus Javanese, Sundanese, and Minangkabau. Their experiments show that Indonesian-centric LLMs frequently produce unsafe responses in colloquial and local languages, and fine-tuning on their data reduces unsafe output without hurting downstream performance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The safety and fine-tuning claims rest on a GPT-4o judge that also generated the fine-tuning targets; with limited human validation and a ~30% false-negative rate on the unsafe class, the key numbers in Tables 2 and 3 are not yet established.","rationale":"The reader's weakest assumption already identifies the single-judge problem. My reading converges on that point and adds a specific mechanism: because GPT-4o both generates the safe training responses and evaluates the fine-tuned outputs, the Table 3 improvement is vulnerable to evaluator self-preference, not just to imperfect cultural understanding. The paper deserves credit for reporting the Appendix G human check and for explicitly listing this limitation, but the check is too small to support the claims as stated: it covers only formal and colloquial Indonesian, has a 30% false-negative rate on unsafe cases, and says nothing about the three local languages or the before/after fine-tuning comparison. The downstream benchmark results (Table 4) are not affected by this concern, and the dataset-construction pipeline is human-reviewed, so the appropriate disposition is unchanged: conditional acceptance with the required remedy being broader human validation and data release. I am not proposing rejection: the dataset may well be valuable, but the empirical claims about local-language safety and fine-tuning effectiveness are not yet settled by the evidence in the paper.","tokens_in":25038,"tokens_out":6598,"duration_ms":68889,"concrete_test":"Extend the Appendix G human-vs-GPT-4o comparison to (1) a stratified sample of 100 responses per local-language variant (Javanese, Sundanese, Minangkabau), and (2) 100 Sailor2 responses before and after fine-tuning, annotated by native speakers using the same evaluation questions. Recompute Tables 2 and 3 with human labels only. If the human-verified unsafe counts for local languages differ from GPT-4o by more than ~20%, or if the before/after reduction in Table 3 no longer reaches significance under McNemar's test, the safety-improvement claim must be revised. Adding an independent second judge (e.g., Claude) would also show whether the result is specific to GPT-4o.","verdict_should_be":"UNCHANGED","load_bearing_attack":"All of the paper's safety numbers are produced by GPT-4o answering binary evaluation questions (Section 5.2, Figure 8). The only human check (Appendix G) covers 200 formal/colloquial prompts and shows a failure to flag 9 of 30 human-flagged unsafe responses (4/17 formal, 5/13 colloquial)—a 30% false-negative rate on the unsafe class. No human check is reported for Javanese, Sundanese, or Minangkabau, where the region-specific risk categories are most dependent on cultural nuance. The fine-tuning result is additionally exposed to circularity: the safe responses in IndoSafety-Train were generated by GPT-4o (Section 4.2, Figure 6), and the same model then judges whether Sailor2's post-tuning outputs are safe. The large reductions in Table 3 (e.g., risk area VI: 53 to 5 formal, 57 to 11 colloquial) may therefore reflect the judge recognizing GPT-4o-style refusal phrasing rather than a verified reduction in culturally unsafe content. The Limitations section explicitly acknowledges the single-judge risk, but an acknowledgment does not validate the reported effect sizes; the load-bearing parts of the central claim still depend on an unvalidated evaluator for exactly the languages and comparisons that matter most.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IndoSafety, a human-verified safety evaluation dataset for Indonesian and three local languages (Javanese, Sundanese, Minangkabau) as well as formal and colloquial Indonesian. It extends the taxonomy of Wang et al. (2024c) with a region-specific risk area (VI) covering 19 harm types, builds a 2,514-prompt evaluation set and a 2,500-prompt parallel test set, and uses the remaining 2,014 prompts as a safety-alignment training set. The authors report unsafe-response rates for ten LLMs across variants (Tables 2 and 10), analyze behavior by risk area and prompt type, and fine-tune Sailor2 with LoRA on IndoSafety-Train, reporting large judged safety improvements (Table 3) with little loss on Indonesian benchmarks (Table 4).","tokens_in":25301,"tokens_out":7237,"duration_ms":66855,"significance":"If the evaluator reliability concerns are resolved, this is a genuinely useful resource: it is the first Indonesian safety evaluation dataset with human-verified prompts across five language varieties, it proposes a culturally grounded taxonomy, and it ships both evaluation and training splits together with a multi-model comparison. The paper is honest in its limitations and follows established evaluation practice, but the central quantitative claims currently rest on a single LLM judge whose agreement with humans is checked only on 200 formal/colloquial prompts. Because the dataset itself is independently human-verified and the taxonomy is well motivated, the contribution does not collapse; it needs stronger validation of the automatic judge and of the fine-tuning effect before the headline numbers can be taken at face value.","major_comments":[{"comment":"All unsafe-response rates in Tables 2 and 3 and Figure 10 are produced by GPT-4o answering binary evaluation questions (Section 5.2, Figure 8). The only human comparison (Appendix G) covers 100 formal and 100 colloquial prompts and shows that GPT-4o rated as safe 9 of the 30 responses human annotators flagged unsafe (4/17 formal, 5/13 colloquial), a 30% false-negative rate on the unsafe class, and no human check is reported for Javanese, Sundanese, or Minangkabau. These are exactly the variants where the new region-specific category VI is most dependent on cultural nuance. The reported effect sizes, including the regional generalization claim, are therefore not yet established. I recommend a stratified human validation sample for the local-language variants, with particular attention to risk area VI, and agreement metrics appropriate for binary decisions (e.g., Cohen's kappa and class-wise recall), not only Pearson correlation.","section":"Section 5.2 and Appendix G"},{"comment":"The fine-tuning demonstration is exposed to a circularity concern: the safe responses in IndoSafety-Train were generated by GPT-4o (Section 4.2, Figure 6), and the harmfulness of the fine-tuned model's outputs is judged by GPT-4o (Section 5.2, Figure 8). The large reductions in Table 3 (e.g., risk area VI from 53/57/60/54 to 5/11/15/15 in colloquial/formal/Javanese/Sundanese) may therefore reflect the judge recognizing GPT-4o-style refusal phrasing rather than a verified reduction in culturally unsafe content. The Limitations section acknowledges the single-judge risk, but an acknowledgment does not validate the effect sizes. I ask for an independent evaluation of a sample of pre- and post-tuning outputs, preferably human-annotated or judged by a model not involved in generating the training targets, with the judge blind to whether each response is from the base or fine-tuned model.","section":"Section 4.2, Section 5.2, Table 3"}],"minor_comments":[{"comment":"The confusion matrices in Figures 11 and 12 are difficult to read because the row and column labels are ambiguous; please draw them as standard confusion matrices with human annotation on one axis and GPT-4o prediction on the other, and report the cell counts in the caption.","section":"Appendix G"},{"comment":"The statement that all differences are significant at alpha = 0.05 under McNemar's test is not accompanied by p-values; given the number of comparisons across variants and risk areas, please report exact p-values or corrected q-values to make the significance claim auditable.","section":"Table 3"},{"comment":"The language-identification accuracy of 97% and prompt-type classification accuracy of 96% are reported only as aggregate numbers based on 100 manual samples; a per-language breakdown would be more informative, especially for Minangkabau, where Figure 4 already shows mixed responses.","section":"Section 6.2"},{"comment":"The translation quality, fluency, and relevance scores in Figure 3 are reported without confidence intervals or the number of annotators; please state whether the values are percentages, how ties were resolved, and whether the 100 samples were scored by a single annotator.","section":"Section 4.1.1 and Figure 3"},{"comment":"The exclusion of Cendol and Komodo because they 'performed poorly' is reported only in a footnote; specifying the failure mode (e.g., empty responses or language mismatch) would help readers interpret model coverage.","section":"Section 5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a well-motivated resource paper whose main gap is methodological validation rather than novelty. If the authors add human validation for the local-language variants and a judge-independent check of the fine-tuning result, I would support acceptance. The paper fits the scope of the venue and the related work is appropriately cited."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper to know about. It delivers the first human-verified safety evaluation benchmark for Indonesian, colloquial Indonesian, and three local languages—Javanese, Sundanese, Minangkabau—with a culturally grounded taxonomy that includes Pancasila misuse, supernatural claims, and ethnic/religious sensitivities. The dataset construction is careful: native-speaker verification, localization of entities and units, and parallel prompts across five variants. The paper also evaluates 10 models and shows that local languages are riskier, which is a plausible and important finding. That part I trust.\n\nThe soft spots are all in the evaluation methodology. The only judge is GPT-4o, and the human agreement check covers 200 formal/colloquial prompts: Pearson r~0.71, with 9 of 30 human-flagged unsafe responses missed. That is a 30% false-negative rate on the unsafe class, and there is no human check at all for Javanese, Sundanese, or Minangkabau. So the absolute unsafe rates in Table 2 are likely underestimated, and the cross-language comparisons are noisy. The authors acknowledge the single-judge risk in Limitations, but the acknowledgment does not fix the numbers.\n\nThe fine-tuning result has a bigger problem. The safe responses in IndoSafety-Train were generated by GPT-4o, and the same model then judges whether the fine-tuned model's outputs are safe. The large improvements in Table 3 (e.g., risk area VI from 53-57 down to 5-11) may partly reflect the judge recognizing GPT-4o-style refusal phrasing rather than a verified reduction in culturally unsafe content. This is also acknowledged, but the claim \"significantly improves safety\" is not yet established. A human evaluation on a sample of pre/post outputs would settle it.\n\nAlso, the dataset is not released, and there is no link in the paper. For a resource paper, that is a real problem.\n\nNet: the dataset and taxonomy are worth having, and the paper deserves a serious referee. But the evaluation story needs major work: release the data, validate the judge on local languages, and add a human-annotated sample to the fine-tuning evaluation. If those are fixed, this could become a useful benchmark for Indonesian safety. I would send it to review with that expectation.","headline":"IndoSafety fills a real gap in Indonesian safety evaluation, and the dataset/taxonomy are solid contributions, but the GPT-4o-only judge and fine-tuning circularity mean the headline numbers should be treated as provisional.","tokens_in":25829,"tokens_out":2127,"would_cite":true,"duration_ms":20444,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning on local safety data cuts unsafe AI outputs in Indonesian","keywords":["LLM safety","culturally grounded safety","Indonesian languages","safety evaluation dataset","Javanese","Sundanese","Minangkabau","safety fine-tuning"],"falsifier":"Take a random sample of the Javanese, Sundanese, and Minangkabau response sets, have native-speaker annotators apply the paper's own per-harm question sets, and compare their labels with GPT-4o's; if agreement on these local languages falls to or below the 0.71–0.72 Pearson correlation seen for Indonesian, the reported unsafe rates and the measured fine-tuning improvement would need recomputation.","tokens_in":24855,"feed_emoji":"🛡️","tokens_out":7078,"duration_ms":59072,"temperature":0.7,"pith_summary":"IndoSafety is a human-verified safety evaluation dataset built for Indonesia's linguistic landscape: formal and colloquial Indonesian plus Javanese, Sundanese, and Minangkabau. The paper argues that translated English safety benchmarks miss what is actually dangerous in this context — ethnic stereotypes, misrepresentation of traditional practices, misinterpretation of Pancasila, regional separatism advocacy, and treating the supernatural as fact — and that this blind spot matters because Indonesia is populous, culturally diverse, and rapidly adopting AI. Evaluating ten LLMs, it finds that Indonesian-centric and regional models produce unsafe responses at substantial rates, worst in colloquial and local-language settings. It then shows that LoRA fine-tuning an open-weight model on the dataset's training split cuts unsafe responses across all risk areas, including Javanese and Sundanese that were not used in tuning, while leaving downstream task performance essentially unchanged. If right, the work makes a concrete case that LLM safety must be evaluated and trained in local cultural terms, not merely translated.","feed_headline":"Fine-tuning on local safety data cuts unsafe AI outputs in Indonesian","feed_subtitle":"Human-verified benchmark across five languages shows local dialects expose safety gaps that tuning closes.","key_machinery":"The carrying mechanism is a fine-grained safety taxonomy: risk areas I–V (discrimination and toxicity, human–chatbot interaction harms, information hazards, malicious uses, misinformation harms) are adopted from the Do-Not-Answer framework, and a new risk area VI, 'region-specific sensitivities', adds eight culturally grounded harm types — ethnicities and cultural practices, historical controversies, Indonesian entities, Pancasila misinterpretation, regional separatism advocacy, religions and beliefs, and supernatural claims. Around this taxonomy the paper builds a three-part dataset: IndoSafety-Eval-1 (2,514 prompts), a parallel test set IndoSafety-Eval-2 (500 prompts in each of five variants), and IndoSafety-Train (2,014 prompt–response pairs with GPT-4o-generated safe answers). Harmfulness is scored by a GPT-4o judge answering per-risk-area binary question sets in Indonesian, adapted from a prior Chinese safeguard evaluation for areas I–V and newly written for area VI. The fine-tuning demonstration uses LoRA on the 8B Sailor2 model for one epoch.","core_discovery":"The paper's central claim is that culturally grounded safety for LLMs cannot be supplied by translated English datasets, and that a purpose-built, human-verified resource changes both measurement and behavior. With IndoSafety, the authors introduce the first safety evaluation dataset tailored to the Indonesian context, spanning five language varieties and a 19-category taxonomy. Their measurements show that existing Indonesian-centric LLMs often generate unsafe outputs — Sailor2 at 36.7% unsafe on the first evaluation set and 32–40% across colloquial, Javanese, and Sundanese variants, with region-specific sensitivities among the most frequent failure areas. The intervention result is the load-bearing outcome: fine-tuning Sailor2 with 2,014 safety prompt–response pairs in formal Indonesian reduced unsafe responses in formal and colloquial Indonesian, Javanese, and Sundanese, with all differences significant at $\\alpha = 0.05$ by McNemar's test, while 3-shot accuracy on six Indonesian benchmarks dropped by at most a fraction of a point. The authors read this as evidence that safety alignment in a high-resource standardized variant can generalize to related low-resource languages.","pith_inferences":["Because human agreement with the GPT-4o judge was measured only for formal and colloquial Indonesian and fell at a Pearson correlation near 0.71, the reported unsafe rates for Javanese, Sundanese, and Minangkabau rest on an unvalidated judge; a native-speaker recoding of those 1,500 responses would tell whether the ranking of models survives.","The fine-tuning gains may partly reflect the evaluator's own preferences rather than an absolute safety improvement; testing the same before/after models with a second judge or with human annotators would separate the two.","If the transfer result generalizes, it offers a recipe for other low-resource language clusters: align once in the standardized variety, then evaluate in the dialects — but only after verifying that the evaluator, not just the model, is fluent in those dialects.","The taxonomy's normative choices (e.g., supernatural claims must be labeled unproven, separatist advocacy must be refused) invite an annotator-disagreement study, since reasonable native speakers may differ on some category boundaries."],"forward_implications":["Safety evaluations that rely on translated English prompts will understate risk in Indonesian colloquial and local-language use; a culturally grounded taxonomy is needed to see those failures.","The dataset works as both a benchmark (the two evaluation sets) and a training resource (IndoSafety-Train), so it can be reused for measurement and alignment in one package.","Safety fine-tuning on formal Indonesian data transfers to related low-resource varieties such as Javanese and Sundanese, suggesting a cheaper alignment path than collecting safety data in every language.","Open-weight regional models below 10B parameters can be made substantially safer with a one-epoch LoRA run without meaningful loss on downstream Indonesian benchmarks.","Models that look safe in formal Indonesian can still fail on interrogative prompts and on region-specific sensitivity categories, so deployment screening should include those conditions."],"supporting_citations":[{"why":"Supplies the base five-category safety taxonomy and the Do-Not-Answer prompts that are translated, localized, and merged into IndoSafety-Eval-1.","marker":"Wang et al. (2024c)"},{"why":"Supplies the binary question-set evaluation strategy and the general-safety evaluation questions adapted for Indonesian.","marker":"Wang et al. (2024d)"},{"why":"Supplies additional safety subsets (I-MaliciousInstructions, I-CoNa, Q-Harm) and the precedent that LoRA fine-tuning on a few hundred safety examples improves safety.","marker":"Bianchi et al. (2024)"},{"why":"The QwQ-32B model used for few-shot augmentation of the Indonesian-specific prompts.","marker":"Qwen Team (2025)"},{"why":"Provides GPT-4o, the model used as the automatic safety judge, the colloquial variant translator, and the generator of safe responses for IndoSafety-Train.","marker":"OpenAI et al. (2024)"},{"why":"Provides LoRA, the parameter-efficient fine-tuning method used in the safety-tuning experiment.","marker":"Hu et al. (2021)"},{"why":"Precedent for extending the safety taxonomy with culturally grounded categories in another language context, and cited in the limitations comparison.","marker":"Ashraf et al. (2025)"}],"fun_headline_variants":["Indonesian LLMs get safer when tuned on culturally grounded data","Fine-tuning on IndoSafety reduces unsafe responses across dialects","Human-verified safety data improves Indonesian LLMs without performance loss","IndoSafety benchmark exposes gaps and shows tuning closes them"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Every safety number in the paper — the model rankings and the fine-tuning improvement — comes from a single automatic judge, GPT-4o, whose agreement with humans was checked on only 200 formal and colloquial Indonesian prompts and never on Javanese, Sundanese, or Minangkabau, so its judgment of culturally specific content is assumed rather than shown.","fun_headline_variants_meta":{"raw":{"variants":["Indonesian LLMs get safer when tuned on culturally grounded data","Fine-tuning on IndoSafety reduces unsafe responses across dialects","Human-verified safety data improves Indonesian LLMs without performance loss","IndoSafety benchmark exposes gaps and shows tuning closes them"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00036,"raw_usage":{"total_tokens":1948,"prompt_tokens":948,"completion_tokens":1000,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":932}},"tokens_in":564,"tokens_out":1000,"duration_ms":8709,"temperature":1.0,"reasoning_tokens":932,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:21:01.138299+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the Javanese, Sundanese, and Minangkabau response sets, have native-speaker annotators apply the paper's own per-harm question sets, and compare their labels with GPT-4o's; if agreement on these local languages falls to or below the 0.71–0.72 Pearson correlation seen for Indonesian, the reported unsafe rates and the measured fine-tuning improvement would need recomputation.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Precedent for extending the safety taxonomy with culturally grounded categories in another language context, and cited in the limitations comparison."}],"review_version":1}