{"id":"0bb21c3b-ca22-4017-bc00-24def3bc282d","arxiv_id":"2507.11502","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"A DeepSeek-based model fine-tuned for Hong Kong outperforms general models on Hong Kong benchmarks, but most of those benchmarks are self-authored and unreleased.","lead":"HKGAI-V1 is a 685-billion parameter language model built by fine-tuning DeepSeek to serve Hong Kong's languages, laws, and values, with support for Cantonese, Mandarin, and English. The paper claims it beats general models on Hong Kong-specific tasks, but the evidence relies heavily on benchmarks the authors wrote themselves and did not release.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Self-authored, unreleased benchmarks—with no contamination analysis and no independent scoring—carry the entire claim of Hong Kong-specific superiority; until a held-out parallel evaluation confirms the margins, the central result is not established.","rationale":"The central claim is an empirical superiority result, so the validity of the evaluation is the load-bearing assumption. I checked whether any independent or public evidence could carry that weight. Beaver-zh-hk and Flames are public, but HKGAI-V1's scores on them are still self-reported through GPT-4o scoring without human validation or error bars; the largest claimed advantages (HKMMLU, SafeLawBench, NaVAB, Adversarial HK Value Bench, Table 7) come from instruments the authors control and do not release. The paper's own §6.1 concedes the subjectivity of 'Hong Kong values' and the annotator sample limits, which further weakens the assumption that the benchmarks are representative. My concern is not that the model is useless or that the authors misbehaved; it is that the evidence cannot distinguish generalization from leakage or rubric-fitting. The reader's weakest assumption identifies the same point, so I agree. A held-out parallel benchmark is the decisive check because it preserves the intended zero-shot setting while removing the possibility that the test items were seen during fine-tuning.","tokens_in":18639,"tokens_out":6735,"duration_ms":78344,"concrete_test":"Compile a held-out parallel HKMMLU—same subject mix, item format, and Traditional Chinese register—from Hong Kong public sources dated after HKGAI-V1's training-data cutoff, and have an independent lab administer it zero-shot to HKGAI-V1, DeepSeek-V3, and GPT-4o. If the 4.8-point lead over DeepSeek-V3 on the original HKMMLU collapses on the never-seen items, the SOTA claim is a benchmark artifact; if it persists, contamination is not the sole explanation. In the same run, report per-category accuracy and a bootstrap confidence interval.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's empirical core is §4.4–§4.5 and §5.3, where HKGAI-V1 is declared superior on Hong Kong-specific tasks. For that claim to hold, the instruments must be valid measures of regional value alignment and must not overlap with the data used to fine-tune the model. Neither condition is currently supported. HKMMLU, SafeLawBench, NaVAB, and the Adversarial HK Value Bench are self-authored and unreleased; the 100-question sensitive-query evaluation in Table 7 is described only by aggregate percentages, with no item examples, no inter-annotator agreement, and no blinding. Section 4.3 states that the fine-tuning pipeline constructed a 'local Hong Kong-specific Q-A dataset' and generated preference data from it, so the boundary between training distribution and test distribution is not established. Section 4.5's claim that the model's 'pre-training corpus' is 'extensively populated' with Hong Kong-relevant text is also in tension with §6.3's statement that HKGAI-V1 is only a full-parameter fine-tuned version of DeepSeek. The reported margins (HKMMLU 81.4 vs 76.6; Beaver-zh-hk 88.95 vs 70.41; Flames 68.06 vs 30.12) could therefore reflect benchmark leakage or rubric design rather than genuine regional alignment. This is not an accusation of fraud; it is an evidential gap that the paper, by keeping model and benchmarks closed, leaves unresolvable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports HKGAI-V1, a 685-billion-parameter model obtained by full-parameter fine-tuning of DeepSeek, combined with RLHF, language-feedback correction via a custom HKValue-Aligner, and a modular retrieval-augmented generation (RAG) system, with the stated goal of regional value alignment for Hong Kong. The authors claim that HKGAI-V1 outperforms general-purpose models on Hong Kong-specific knowledge and safety benchmarks (HKMMLU, SafeLawBench, NaVAB, the Adversarial HK Value Benchmark, and a sensitive-question test set) while preserving general knowledge and reasoning (MMLU, AGI-Eval). The paper also discusses governance, deployment in Hong Kong government services, and a roadmap for a future HKGAI-V2.","tokens_in":18959,"tokens_out":4586,"duration_ms":51432,"significance":"If the empirical claims were fully substantiated, the work would be a valuable case study in region-specific alignment and sovereign LLM development, particularly for multilingual settings with Cantonese, Mandarin, and English, and for legal safety in a distinct socio-legal context. The paper is candid about several limitations in Sections 6.1 and 6.3, including annotator subjectivity, a compliance-robustness trade-off, and the model's dependence on DeepSeek. However, the central claim of superiority on Hong Kong-specific tasks rests on proprietary, unreleased benchmarks with no contamination analysis, no independent scoring, and no statistical uncertainty quantification. As presented, the contribution is primarily a system-building narrative; the measurement of regional alignment superiority is not yet established.","major_comments":[{"comment":"There is a load-bearing internal inconsistency between the HKMMLU discussion and the stated model provenance. Section 4.5 attributes the zero-shot HKMMLU result to a pre-training corpus 'extensively populated' with Hong Kong-relevant Traditional Chinese text, but Section 6.3 explicitly states that HKGAI-V1 is a full-parameter fine-tuned version of DeepSeek and that its core capabilities are predetermined by the original training. Since the authors did not pre-train the model, the zero-shot HKMMLU performance must be attributed to fine-tuning data, and the paper does not establish that the fine-tuning data and the HKMMLU test items are disjoint. This affects the central claim of regional superiority and must be resolved by clarifying data provenance and providing contamination analysis.","section":"§4.5 vs §6.3"},{"comment":"The strongest claimed results come from self-authored, unreleased benchmarks: HKMMLU, SafeLawBench, NaVAB, and the Adversarial HK Value Benchmark. The paper provides no contamination checks, no item-level examples, no inter-annotator agreement, and no external or blinded scoring. Given that Section 4.3 describes constructing a local Hong Kong-specific Q-A dataset and generating preference data from it, the boundary between training distribution and test distribution is not demonstrated. The margins reported (e.g., HKMMLU 81.4 vs 76.6; Table 5 HK-sensitive 79% vs ChatGPT 10.7%) cannot be interpreted as genuine regional alignment until the benchmarks are released or a held-out external evaluation with contamination control is provided.","section":"§4.5, §5.3"},{"comment":"The sensitive-question evaluation is based on a 100-question test set with no item examples, no scoring rubric details, no blinding procedure, and no inter-annotator reliability. Several reported metrics are non-quantitative labels ('Yes', 'Partially', 'No'), and claims such as 0% refusal and 100% positive/neutral responses are presented without confidence intervals or significance tests. Similarly, Table 5 reports percentages for the Adversarial HK Value Benchmark but does not state the number of questions per module, how the 300 total questions are allocated, or how human evaluators were trained and blinded. These omissions are load-bearing because the paper's central claim depends on these evaluations.","section":"§5.3, Table 7"},{"comment":"All benchmark comparisons are single-point estimates with no error bars, repeated runs, or significance testing. The Beaver-zh-hk evaluation relies on GPT-4o scoring with a four-tier methodology that is not specified, so the robustness of the 88.95 vs 70.41 margin is unclear. Additionally, the Flames improvement (68.06 vs 30.12) is attributed to RAG and retrieval modules, but no ablation distinguishes the contribution of RAG from that of the fine-tuning and alignment pipeline; Table 6 only provides with/without-search language-following rates, not task performance. The paper should either add ablations or temper the causal claim about RAG.","section":"§4.4, §5.3"}],"minor_comments":[{"comment":"In the RLHF objective, the maximization is written over the reward-model parameters 'phi' while the expression contains the policy parameters 'theta'; this should be 'max_theta' for consistency with the preceding sentence.","section":"§4.1"},{"comment":"The Instruction Following row says 'two languages' but then lists four (Simplified Chinese, Traditional Chinese, English, Cantonese); correct the count or split the row.","section":"§5.3, Table 4"},{"comment":"There are several typographical errors, including 'demonstar' (§6.1), 'standtards' (§5.3), 'Multiple-Choise' (Table 2 note), and 'endowered' (Conclusion). A copyediting pass is needed.","section":"Throughout"},{"comment":"References [31] and [66] appear to be the same arXiv paper (arXiv:2504.12911), and the acronym NaVAB is never expanded; add the full name and deduplicate the references.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The reader's reject verdict is understandable given the evidential gap, but I see the gap as repairable in revision if the authors provide a contamination analysis, release or externally validate the benchmarks, add uncertainty quantification, and resolve the provenance inconsistency in §4.5 vs §6.3. If the benchmarks cannot be released or externally validated, the central claim should be downgraded accordingly, and the paper should be reframed as a system description rather than an empirical superiority claim. I also note that the proprietary benchmark strategy may be intentional for domain-specific evaluation, but scientific claims require independent verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's the short version: HKGAI-V1 is a legitimate systems paper—a 685B DeepSeek fine-tune with RLHF/DPO, correction learning, and RAG, deployed to ~20,000 Hong Kong government users. That deployment is real, and the authors are candid that the model is a fine-tune, not a from-scratch sovereign model. The paper is most useful as a case study in regional alignment practice, not as a methodological advance.\n\nWhat it does well: it documents a complete pipeline, including the HKValue-Aligner correction model and a set of region-specific benchmarks (HKMMLU, SafeLawBench, NaVAB, Adversarial HK Value Bench). The limitations sections (§6.1, §6.3) are honest: they acknowledge the subjectivity of 'Hong Kong values', the limited local corpus, and the model's dependence on DeepSeek. That candor counts for something.\n\nThe soft spots are in the evaluation. The central claim—superiority on Hong Kong-specific tasks—rests entirely on self-authored, unreleased benchmarks. There are no error bars, no significance tests, no contamination analysis, and no independent scoring. The 100-question sensitive-query evaluation (Table 7) is reported as aggregate percentages with no examples and no inter-annotator agreement. That's a real evidential gap. There's also an internal tension: §4.5 says the 'pre-training corpus' is extensively populated with Hong Kong text, but §6.3 says HKGAI-V1 is only a full-parameter fine-tuned version of DeepSeek. Those two statements don't reconcile, and the zero-shot framing in §4.5 is misleading given the model was fine-tuned on HK data.\n\nThe reader's rejection is a bit harsh in tone but the substance holds: the paper cannot be verified as a scientific contribution until the benchmarks are released or an external evaluation is done. That said, this isn't fraud—it's an under-supported applied claim. The paper deserves a serious referee, not because the results are convincing, but because the topic (region-specific alignment and its evaluation) is important and the deployment gives it real-world weight. A good referee could push for benchmark release, a contamination check, and a more careful separation of training and test data.\n\nFor a reading group, it's a decent case study on evaluation pitfalls. I wouldn't cite it for the specific numbers, but I might cite it as an example of deployed sovereign AI. Verdict: reject in current form, but engage with it; the revisions are feasible.","headline":"A deployed Hong Kong sovereign-AI system with standard alignment methods and a self-authored benchmark suite; the regional-superiority claim is plausible but not yet established.","tokens_in":19490,"tokens_out":2766,"would_cite":false,"duration_ms":30545,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HKGAI-V1 is a 685-billion-parameter model, fine-tuned from DeepSeek, that outperforms general-purpose models on Hong Kong-specific culturally sensitive queries and regional value benchmarks while keeping general knowledge and reasoning…","keywords":["sovereign AI","large language model","Hong Kong","value alignment","RLHF","retrieval-augmented generation","Cantonese","regional benchmark"],"falsifier":"Hold out an independently authored set of Hong Kong sensitive questions, written after the model's training data was fixed and in colloquial Cantonese, and compare HKGAI-V1 with DeepSeek-R1 under identical conditions with retrieval turned off; if the Beaver-zh-hk gap of 18.54 points collapses, the regional-alignment claim is mostly benchmark-specific.","tokens_in":18422,"feed_emoji":"🇭🇰","tokens_out":6488,"duration_ms":72018,"temperature":0.7,"pith_summary":"This paper sets out to show that a large language model can be made genuinely regional: a 685-billion-parameter model fine-tuned from DeepSeek, HKGAI-V1 is aligned to Hong Kong's Cantonese-Mandarin-English linguistic environment and its local legal, cultural, and ethical norms. The authors claim that this regional alignment improves performance on Hong Kong-specific sensitive and safety questions while leaving general knowledge and reasoning essentially intact, with notable gains on benchmarks that reward local values (88.95 vs 70.41 on Beaver-zh-hk, and 68.06 vs 30.12 on Flames) and a new state of the art on the author-built HKMMLU (81.4%). They also contribute a reusable evaluation tool, the Adversarial HK Value Benchmark, and a RAG architecture for grounded, time-sensitive answers. If the claims hold, they would show that regional digital sovereignty need not require training from scratch, and that a carefully fine-tuned open base model can outperform major general-purpose systems on locally defined criteria.","feed_headline":"Hong Kong's own LLM beats general models on local value tests","feed_subtitle":"A 685B DeepSeek-based model scores 81.4% on Hong Kong MMLU and 88.95 on local safety, without losing general skill.","key_machinery":"The correction-amplified RLHF loop is the load-bearing mechanism. Local annotators produce a question-answer-correction dataset; the correction model HKValue-Aligner learns to rewrite unsafe or misaligned answers into compliant ones; those corrected answers become preference pairs for RLHF with a KL penalty to the base policy. Around this sits a modular RAG pipeline (intent classifier, retriever, memory, tool use, workflow moderation) and the paper's own evaluation instruments, especially the Adversarial HK Value Benchmark, which scores responses to 300 sensitive questions as safe, template refusal, or unsafe.","core_discovery":"The central discovery is that full-parameter fine-tuning of a strong general-purpose base model, guided by local annotator corrections and RLHF, can encode region-specific values without a general-capability tax. HKGAI-V1 beats DeepSeek-R1 on Beaver-zh-hk by 18.54 points and on Flames by 37.94 points, achieves 81.4% on HKMMLU versus 76.6% for DeepSeek-V3, and is rated safe on 79% of the paper's Hong Kong sensitive adversarial questions, compared with ChatGPT's 10.7%. The authors interpret these results as evidence that a sovereign model can control its own safety behavior, language matching, and local knowledge while remaining competitive on MMLU and AGI-Eval.","pith_inferences":["If the benchmark results generalize beyond the paper's own test sets, then value alignment for a region can be treated as a correction-and-preference amplification problem rather than a pre-training problem, which would lower the barrier for other mid-sized jurisdictions.","The comparison on the Adversarial HK Value Bench mixes refusal with safety: ChatGPT's 10.7% 'safe' rate on Hong Kong sensitive questions could partly reflect a different refusal policy rather than a genuine lack of alignment, so the headline gap is likely smaller than it appears.","A natural next experiment is to run the same pipeline with Cantonese-only prompts and code-mixed vernacular text, since the paper's strongest language-following results are for written Chinese and English while oral Cantonese still shows search-dependent variation.","The governance-embedded architecture, which separates trust, platform, model, and data layers, could be reused as a public-sector template even if the model itself remains based on DeepSeek."],"forward_implications":["Hong Kong public-sector deployments can use a locally controlled model that answers sensitive local questions directly, without ideological templates or refusal, while staying compliant with local law.","A region can obtain a sovereign model by fine-tuning an existing open base model on local data, cutting the compute and data costs of training from scratch.","Regional value alignment does not have to trade away general knowledge: MMLU drops by only 0.36 points while regional safety scores rise sharply.","The same evaluation recipe (a local knowledge benchmark, a legal safety benchmark, and an adversarial value benchmark) can be transplanted to other jurisdictions.","The 16.5% unsafe rate on instruction attacks is a known consequence of regional alignment and defines the next target for the V2 roadmap."],"supporting_citations":[{"why":"The base model HKGAI-V1 is full-parameter fine-tuned from DeepSeek, and DeepSeek-R1 is the main general-purpose baseline.","marker":"[17]"},{"why":"The main Hong Kong safety benchmark, built from 29 risk scenarios; HKGAI-V1 scores 88.95 vs 70.41.","marker":"[20]"},{"why":"Chinese value-alignment benchmark on which the model jumps from 30.12 to 68.06.","marker":"[22]"},{"why":"Author-built Hong Kong MMLU benchmark on which HKGAI-V1 sets the reported 81.4% SOTA.","marker":"[7]"},{"why":"Legal-safety benchmark used to measure alignment with local laws.","marker":"[65]"},{"why":"Multi-national value alignment benchmark used to show cross-cultural value alignment.","marker":"[66]"},{"why":"The Aligner correction method behind the local HKValue-Aligner.","marker":"[24]"},{"why":"Language-feedback method used to synthesize preference data for RLHF.","marker":"[29]"},{"why":"The standard RLHF objective with a Bradley-Terry reward model and KL penalty.","marker":"[40]"},{"why":"The retrieval-augmented generation approach that grounds the RAG module.","marker":"[34]"}],"fun_headline_variants":["HKGAI-V1: Hong Kong LLM beats general models on local values","Sovereign HK model beats generalists on Cantonese ethics tests","HKGAI-V1: full-parameter tuning secures HK values, no skill loss","Hong Kong's own AI: 79% safe on sensitive local questions vs 10.7%","HKGAI-V1: region-specific alignment without a general capability tax"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's headline results rest on benchmarks the authors built themselves, so the central assumption is that those benchmarks genuinely represent Hong Kong values and are not contaminated by the fine-tuning data.","fun_headline_variants_meta":{"raw":{"variants":["HKGAI-V1: Hong Kong LLM beats general models on local values","Sovereign HK model beats generalists on Cantonese ethics tests","HKGAI-V1: full-parameter tuning secures HK values, no skill loss","Hong Kong's own AI: 79% safe on sensitive local questions vs 10.7%","HKGAI-V1: region-specific alignment without a general capability tax"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000897,"raw_usage":{"total_tokens":3885,"prompt_tokens":987,"completion_tokens":2898,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":2791}},"tokens_in":603,"tokens_out":2898,"duration_ms":22806,"temperature":1.0,"reasoning_tokens":2791,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:32:39.652499+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Hold out an independently authored set of Hong Kong sensitive questions, written after the model's training data was fixed and in colloquial Cantonese, and compare HKGAI-V1 with DeepSeek-R1 under identical conditions with retrieval turned off; if the Beaver-zh-hk gap of 18.54 points collapses, the regional-alignment claim is mostly benchmark-specific.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The main Hong Kong safety benchmark, built from 29 risk scenarios; HKGAI-V1 scores 88.95 vs 70.41."},{"cited_title":"Measuring Hong Kong Massive Multi-Task Language Understanding","cited_arxiv_id":"2505.02177","evidence_quote":"Author-built Hong Kong MMLU benchmark on which HKGAI-V1 sets the reported 81.4% SOTA."},{"cited_title":"& Gu o, Y","cited_arxiv_id":null,"evidence_quote":"Legal-safety benchmark used to measure alignment with local laws."},{"cited_title":"Benchmarking Multi-National Value Alignment for Large Language Models","cited_arxiv_id":"2504.12911","evidence_quote":"Multi-national value alignment benchmark used to show cross-cultural value alignment."},{"cited_title":"& Yang, Y","cited_arxiv_id":null,"evidence_quote":"The Aligner correction method behind the local HKValue-Aligner."},{"cited_title":"& Lowe, R","cited_arxiv_id":null,"evidence_quote":"The standard RLHF objective with a Bradley-Terry reward model and KL penalty."},{"cited_title":"& Kiela, D","cited_arxiv_id":null,"evidence_quote":"The retrieval-augmented generation approach that grounds the RAG module."}],"review_version":1}