{"id":"e421d1f9-a791-4765-b2cd-4ac7187aa5f2","arxiv_id":"2607.20410","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A survey-derived Sri Lankan value alignment suite (LKValues) with 150k instruction instances and a 1k benchmark improves Qwen-family LLMs' Sri Lankan value judgment in Sinhala and English.","lead":"LKValues introduces a survey-grounded Sinhala-English dataset suite for aligning LLMs with Sri Lankan societal values, including a 150k-instruction corpus and a 1k-item benchmark. Fine-tuning on this data improves Qwen-family models' value-sensitive judgment in English and Sinhala, though gains do not transfer to Aya-Expanse-8B.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'majority-endorsed' value inventory is statistically fragile: at n≈205 and 85% MOE, several retained values (Equality 52.5%, Justice 52.7%) have CIs spanning 50%, so the survey grounding of LKValues is weaker than claimed.","rationale":"The reader's weakest_assumption correctly identified the survey as the load-bearing premise. I sharpen it: the >50% retention rule is not statistically robust for a non-negligible subset of the 40 values, even under the paper's own 85% confidence choice. This is more specific than a general representativeness critique, and it is testable from the paper's reported numbers. I considered the alternative concern of benchmark-label reliability: Section 4.3 reports agreement for news tagging and scenario generation but not for LKvaluesBench gold labels. That is a real gap, but the fine-tuned models generalize to unseen SinhalaMMLU-derived items, which mitigates pure overfitting. I also considered circularity between LKvaluesIT and LKvaluesBench; the inclusion of human-curated SinhalaMMLU items weakens that objection. The statistical fragility of the value inventory, however, directly undermines both the construct and the criterion validity of the resource, and it is acknowledged only partially in the Limitations. The proposed 95% CI check would settle whether the borderline values actually clear the majority bar; if they do not, the paper should either use a stricter threshold or re-scope claims as 'survey-derived' rather than 'majority-endorsed.' This does not change the reader's conditional verdict: the paper is promising but needs this robustness check before acceptance.","tokens_in":44063,"tokens_out":8825,"duration_ms":73884,"concrete_test":"Using the reported per-value endorsement proportions and item-wise n (Figure 3), compute 95% confidence intervals for each of the 40 retained values (normal approximation with FPC, or exact binomial). Count how many have a lower bound ≤50%. Then re-run the main LKvaluesBench evaluation excluding items whose primary value fails this majority test; if the Qwen3.5-4B-FullSFT-LKV gains (0.610→0.863 AC, 0.491→0.801 SI AC) attenuate or reverse, the central claim depends on borderline values.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The resource suite's central claim rests on the 40 'majority-endorsed' values identified in Section 3 and used to build LKvaluesIT and LKvaluesBench. The endorsement threshold is >50%, but with n≈205 and an 85% confidence MOE of ~±5 pp (Appendix A.7, Figure 3), values near the threshold are not statistically distinguishable from 50%. For instance, Equality is 52.48% ±5.06 (lower bound 47.4%), Justice 52.66% ±5.14, and Belonging 55.94% ±5.03. At the customary 95% confidence (z=1.96), these intervals widen to roughly ±6.9 pp, making even more of the 40 values fail the majority test. The Limitations section concedes the sample is small and subgroup-imbalanced (e.g., Burgher n=3), but the main text still calls the inventory 'majority-endorsed.' Because the benchmark gold labels and training supervision are both derived from this inventory, an unsupported threshold propagates into the measured improvements: the headline gain for Qwen3.5-4B-FullSFT-LKV (0.610→0.863 AC) could partly reflect alignment to values that are not actually endorsed by a majority. This is a statistical robustness issue, not a dispute with the authors' good-faith effort.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents LKValues, a Sinhala-English resource suite for aligning LLMs with Sri Lankan societal values. The authors collect a trilingual survey from 205 respondents, derive 40 values endorsed by more than 50% of the sample, and use this inventory to construct LKvaluesIT (a 150k-instance instruction corpus from Sri Lankan news) and LKvaluesBench (a 1,000-instance bilingual value-sensitive judgment benchmark). They evaluate proprietary and open-weight models, then fine-tune Qwen3.5-4B, Qwen3.5-9B, and Aya-Expanse-8B using LKvaluesIT mixed with general Sinhala instruction data. The headline claim is that LKValues fine-tuning substantially improves Qwen-family models, e.g., Qwen3.5-4B-Base improves from 0.610 AC / 0.491 SI AC / 18.20% invalid outputs to 0.863 AC / 0.801 SI AC / 0.00% invalid outputs with full SFT (Table 1), while gains remain model-family dependent, with Aya-Expanse-8B not benefiting under the same one-epoch LoRA recipe.","tokens_in":44472,"tokens_out":6224,"duration_ms":51068,"significance":"If the resource is taken at face value, LKValues is a useful contribution to low-resource, country-specific value alignment: the dataset is publicly released, grounded in Sri Lankan news, bilingual, and evaluated through a broad protocol with human validation and reported inter-annotator agreement for several stages. The negative result for Aya-Expanse is also informative, as it documents that a single fine-tuning recipe does not transfer across model families. However, the central claim is currently supported only under the authors' own operationalization of Sri Lankan values. The survey sample is small, convenience-based, and statistically fragile near the 50% threshold, and both the training data and the evaluation benchmark are derived from the same value inventory. The resource is valuable as a replicable pipeline and a first benchmark, but the stronger generalization claims in the abstract and conclusion need additional external validation before they can be accepted at face value.","major_comments":[{"comment":"The 'majority-endorsed' label is not statistically supported at the reported precision. With n≈205 and an 85% confidence MOE of about ±5 pp, values such as Equality (52.48±5.06), Justice (52.66±5.14), and Belonging (55.94±5.03) have intervals that cross 50%; at the conventional 95% level the intervals widen further. Because the same 40-value inventory is used to construct LKvaluesIT and LKvaluesBench (§4), both the training supervision and the benchmark gold labels inherit this uncertainty. The paper should either use an uncertainty-aware criterion (e.g., requiring the lower confidence bound to exceed 50%), report sensitivity analyses removing values whose CI crosses the threshold, or explicitly demote 'majority-endorsed' to 'survey-endorsed under the authors' operationalization.' In addition, Figure 3 is captioned 'all 40 Sri Lankan societal values' but tabulates 51 rows; the retained 4","section":"§3, Appendix A.7, Fig. 3"},{"comment":"There is a circularity between training and evaluation: LKvaluesIT value labels and LKvaluesBench gold labels are both derived from the same 40-value survey inventory, and the benchmark items are created or adapted to fit those values. The reported gains (e.g., Qwen3.5-4B base 0.610 AC to 0.863 AC FullSFT in Table 1) therefore measure the model's ability to reproduce the authors' operationalization of Sri Lankan values, not alignment with an independently established set of societal values. This does not invalidate the resource, but it weakens the generalization claim in the abstract and conclusion. I ask for at least one additional validation: human-model agreement on LKvaluesBench, a held-out set of items not constructed from the inventory, or a comparison against an external value-alignment benchmark.","section":"§1, §4.1, §4.2"},{"comment":"Benchmark gold-label reliability is not quantified. Quality control reports Fleiss' κ=0.81 for value tagging and κ=0.82/0.75 for scenario generation (Section 4.3), but no inter-annotator agreement is reported for the LKvaluesBench A/B/BOTH/0 labels, which are the direct evaluation target. This is especially important because 509 of the 1,000 benchmark items are LLM-generated and only 'human-verified.' Please provide per-item agreement statistics (e.g., κ or adjudication rate), the number of annotators per item, and a clear description of how disagreements were resolved.","section":"§4.2, §4.3"},{"comment":"The reported headline numbers differ across tables without definitional reconciliation. Table 1 reports Qwen3.5-4B-FullSFT-LKV AC=0.863 and Qwen3.5-9B-LoRA-LKV AC=0.736; Table 8 reports the similar-sounding 'Value-align. AC' as 0.868 and 0.717, and Table 11's language averages for Aya-Expanse-8B-LoRA-LKV ((0.730+0.408)/2=0.569) match Table 1 but not Table 8 (0.555). If the appendix metric is computed on a different subset (e.g., retained value categories) or with different averaging, this must be stated explicitly. Without clarification, a reader cannot verify the central quantitative claim.","section":"Table 1, Table 8, Table 11"}],"minor_comments":[{"comment":"Typo: 'The survey designs to capture' should be 'The survey is designed to capture.'","section":"§3"},{"comment":"The citation 'hae (2022)' appears garbled; this likely refers to the World Values Survey (Haerpfer et al., 2020) or similar and should be corrected.","section":"§1"},{"comment":"The citation 'Political compass1 (Röttger et al., 2024)' is confusing: Röttger et al. is a critical methodological paper, not the Political Compass instrument itself. Please clarify how this source was used for question selection.","section":"§1"},{"comment":"Footnote markers such as '3,000 years16' and 'Charter for a Pluralistic Sri Lankan Society17' do not have corresponding footnotes in the visible text; please fix the reference formatting.","section":"Appendix A.2"},{"comment":"Minor presentation issues: Table 3 caption says '40 values' but the table lists 51 rows; Table 15 is very long and could be summarized in the main text with a pointer to the appendix; several model names in the references have inconsistent spacing (e.g., 'Pengyun Zhu 1').","section":"Various tables"}],"recommendation":"major_revision","confidential_remarks":"The paper makes a good-faith contribution to a genuinely under-served area, and the resource could be valuable to the community. My main reservation is that the central empirical claim rests on a statistically fragile survey inventory and a potentially circular evaluation. A revision that adds uncertainty-aware analysis, external validation, and explicit clarification of the value inventory and metric definitions would substantially strengthen the paper. I would be receptive to a revised version along these lines."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"LKValues is a real contribution: the first survey-grounded Sinhala–English value alignment resource for Sri Lanka, with a 150k instruction set, a 1k benchmark, and public release. The pipeline is replicable, the human validation is careful (Fleiss κ 0.81–0.82), and the authors honestly report that the same LoRA recipe fails on Aya-Expanse. The headline result—Qwen3.5-4B-Base going from 0.610 to 0.863 AC with full SFT—is supported by the tables.\n\nThe soft spots are real but not disqualifying. The stress-test note lands: with n=205 and the reported 85% MOE, retained values near the threshold (Equality 52.5%, Justice 52.7%) have confidence intervals that dip below 50%. Calling these 'majority-endorsed' overstates the statistical footing; the paper's own caveat in Section 3 acknowledges this, but the abstract and framing do not. The deeper issue is circularity: LKvaluesIT and LKvaluesBench are both built from the same 40-value inventory, so fine-tuning gains show the model matching the authors' operationalization, not necessarily generalizing to Sri Lankan values as a whole. That is an inherent limitation of this kind of resource, but it should be stated more prominently than it is.\n\nThere are also minor table inconsistencies (e.g., Qwen3.5-4B-LoRA-LKV AC is 0.790 in Table 1 but 0.777 in Table 8) that make precise numbers hard to trust without checking the appendix.\n\nWho should read it: anyone building low-resource cultural alignment resources, or benchmarking multilingual value-sensitive reasoning. It deserves a serious referee; the data are useful and the flaws are fixable in revision. The authors should add a robustness analysis treating the threshold as uncertain and clarifying that the benchmark measures alignment to the survey operationalization.","headline":"A genuinely useful Sinhala-English value alignment resource, but the 'majority-endorsed' inventory is statistically fragile and the training/benchmark circularity narrows what the gains mean.","tokens_in":44952,"tokens_out":3403,"would_cite":true,"duration_ms":27844,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A survey of 205 Sri Lankans yields 40 values that, when fine-tuned into Qwen models, lift Sinhala value-judgment accuracy from 49% to 80% and cut invalid outputs to zero.","keywords":["value alignment","Sri Lankan societal values","Sinhala","low-resource languages","survey-based value identification","instruction fine-tuning","cultural alignment benchmark","LLM evaluation"],"falsifier":"A replication study with a larger, more demographically representative Sri Lankan sample (e.g., stratified by ethnicity, religion, and region) that recomputes endorsement proportions and re-derives the value inventory: if the majority-endorsed set changes substantially, the benchmark and all fine-tuning results built on it would need to be revisited.","tokens_in":44017,"feed_emoji":"🇱🇰","tokens_out":1218,"duration_ms":13782,"temperature":0.7,"pith_summary":"This paper tries to establish that culturally grounded, survey-derived resources can close the value-alignment gap for low-resource languages. It builds LKValues, a bundle of 40 majority-endorsed Sri Lankan societal values, a 150k-instance bilingual instruction corpus, and a 1,000-item evaluation benchmark. The core claim is that fine-tuning on these resources substantially improves Qwen-family models' value-sensitive judgment in both English and Sinhala, reducing invalid outputs and cross-lingual disparities — though gains are model-family dependent. A sympathetic reader would care because it offers a replicable, country-specific alternative to Western-centric alignment, and shows that scale and recency alone don't guarantee culturally aware behavior.","feed_headline":"Survey-derived values lift Qwen Sinhala judgment by 31 points","feed_subtitle":"Fine-tuning on 150k Sri Lankan scenarios cuts invalid outputs to zero and narrows the English–Sinhala gap.","key_machinery":"The central mechanism is the survey-to-benchmark pipeline: 40 majority-endorsed Sri Lankan societal values derived from a trilingual survey (with a >50% endorsement threshold) organize both the instruction dataset LKvaluesIT and the evaluation benchmark LKvaluesBench. The survey values act as the controlling taxonomy — every training instance is tagged with one value, and every benchmark item maps to one primary value. Fine-tuning uses LKvaluesIT mixed with 50K general-purpose Sinhala instruction examples, and evaluation uses a strict A/B/BOTH/0 forced-choice protocol with dual prompt framings (Sri Lankan-specific vs. Universal).","core_discovery":"The paper claims that LKValues — the first survey-grounded resource suite for Sri Lankan value alignment — improves value-sensitive judgment of LLMs in Sinhala and English. From a trilingual survey of 205 respondents, the authors derive 40 majority-endorsed values, then build LKvaluesIT (150k scenario-based instruction instances) and LKvaluesBench (1,000 evaluation instances). Fine-tuning Qwen3.5-4B-Base with full SFT raises its accuracy from 0.610 to 0.863 overall and from 0.491 to 0.801 in Sinhala, while reducing invalid outputs from 18.20% to 0.00%. Gains are not universal: the same one-epoch LoRA recipe does not transfer to Aya-Expanse-8B, whose adapted variant drops on the benchmark. Th","pith_inferences":["The paper's 40-value inventory may be an artifact of its 205-respondent online sample and its >50% endorsement threshold; a more representative or larger sample could yield a different value set, and the benchmark's gold labels would shift accordingly.","A testable extension: the same pipeline could be applied to Tamil, the third official language of Sri Lanka, to see whether a trilingual resource suite narrows the Sinhala–English gap further or reveals value differences across linguistic communities.","The distinction between LKvaluesIT (generation) and LKvaluesBench (judgment) suggests a broader principle: alignment resources may need to separate explanatory fluency from forced-choice judgment, since models can excel at one while failing the other.","The finding that Aya-Expanse drops after LoRA fine-tuning hints that multilingual-strong models may require different training schedules or judgment-formatted supervision, not just more data — a hypothesis the paper leaves open."],"forward_implications":["If the central claim holds, fine-tuning on LKValues can substantially improve Sinhala value-sensitive judgment in Qwen-family models, reducing invalid outputs to zero in the best case.","LKvaluesBench can serve as a diagnostic for low-resource cross-lingual alignment gaps: even strong multilingual models like Command-R-08-2024 show large English–Sinhala accuracy gaps.","The replicable pipeline — survey, value tagging, scenario extraction, bilingual evaluation — could be applied to other countries and low-resource languages.","Smaller fine-tuned models can outperform much larger baselines on country-specific value judgment, suggesting that culturally grounded supervision can compensate for scale.","Prompt wording (Sri Lankan-specific vs. Universal) has smaller effects than language, indicating that Sinhala competence and task-format adherence matter more than framing alone.","The model-family dependence (Qwen gains vs. Aya drop) implies that optimal alignment strategies require model-specific tuning of training duration, learning rate, and adaptation method."],"fun_headline_variants":["Survey-derived values lift Qwen Sinhala judgment by 31 points","First Sinhala value-alignment suite cuts invalid outputs to zero for Qwen","New Sri Lankan value benchmark exposes gaps in larger LLMs","Qwen fine-tuned on 40 Sri Lankan values shrinks English–Sinhala gap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the 40 values derived from a 205-respondent online survey (with a >50% endorsement threshold) are a valid operationalization of Sri Lankan societal values, and that LKvaluesBench items correctly measure alignment to those values.","fun_headline_variants_meta":{"raw":{"variants":["Survey-derived values lift Qwen Sinhala judgment by 31 points","First Sinhala value-alignment suite cuts invalid outputs to zero for Qwen","New Sri Lankan value benchmark exposes gaps in larger LLMs","Qwen fine-tuned on 40 Sri Lankan values shrinks English–Sinhala gap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1407,"prompt_tokens":884,"completion_tokens":523,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":440}},"tokens_in":628,"tokens_out":523,"duration_ms":5036,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T09:53:35.506182+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication study with a larger, more demographically representative Sri Lankan sample (e.g., stratified by ethnicity, religion, and region) that recomputes endorsement proportions and re-derives the value inventory: if the majority-endorsed set changes substantially, the benchmark and all fine-tuning results built on it would need to be revisited.","supporting_citations":[],"review_version":1}