{"id":"9fd0015b-4b1a-48ed-998a-6b28446ba34d","arxiv_id":"2601.12696","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"UbuntuGuard is the first culturally-grounded policy benchmark for African languages, derived from expert-authored queries, showing that English-centric safety evaluations overestimate real-world multilingual performance.","lead":"The paper introduces UbuntuGuard, the first policy-based safety benchmark for African languages built from adversarial queries by 155 local domain experts. A smart generalist should read it to see why current Western-centric AI safety tools may fail in non-English contexts and what is needed for more equitable guardian models.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Expert-authored queries lack reported validation for representing diverse African cultural norms and harm scenarios","rationale":"This matches the reader's weakest_assumption exactly and explains the UNVERDICTED verdict: without methodological details on query validity (absent from the abstract and not verifiable here), the headline findings cannot be confidently assessed. The proposed test is a direct, falsifiable check on the load-bearing step.","tokens_in":1710,"tokens_out":337,"duration_ms":48007,"concrete_test":"Sample 50 queries from the released dataset; recruit 20-30 independent native speakers per major language group (stratified by region and demographics) to rate each query on cultural authenticity and realism of the harm scenario using a 5-point Likert scale plus open comments; compute Fleiss' kappa and mean relevance scores—if mean relevance < 3.5 or kappa < 0.4, the cultural-grounding assumption is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that UbuntuGuard reveals English-centric benchmarks overestimate multilingual safety and that dynamic models fail to localize African contexts—depends on the 155 domain experts' adversarial queries accurately capturing culturally grounded risk signals and local norms. The abstract states these queries are used to derive context-specific policies and reference responses but provides no information on expert selection (e.g., geographic or linguistic representation across Africa's 2000+ languages), query elicitation protocol, inter-expert agreement, or any external validation against community input. If the queries primarily reflect the experts' individual or professional perspectives rather than broader sociocultural consensus, the benchmark's grounding is compromised and performance gaps may not generalize.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces UbuntuGuard, the first policy-based safety benchmark for African languages constructed from adversarial queries authored by 155 domain experts across fields including healthcare. From these queries the authors derive context-specific safety policies and reference responses that capture culturally grounded risk signals. They evaluate 15 models (seven general-purpose LLMs and eight guardian models in static, dynamic, and multilingual variants) and report that English-centric benchmarks overestimate real-world multilingual safety, cross-lingual transfer provides only partial coverage, and dynamic models still struggle to localize African-language contexts.","tokens_in":1844,"tokens_out":555,"duration_ms":31461,"significance":"If the expert queries and evaluation protocols are shown to be valid, the work would be a valuable contribution to equitable AI safety research by filling a gap in low-resource, culturally grounded benchmarks. The explicit comparison of static, dynamic, and multilingual guardian variants and the use of expert-authored adversarial queries are positive features that could guide future policy-aligned model development.","major_comments":[{"comment":"§3 (Benchmark Construction): The description of the 155 domain experts' query authoring process provides no information on expert selection criteria, geographic or linguistic representation across Africa's 2000+ languages, elicitation protocol, inter-expert agreement statistics, or any external validation against community input. This is load-bearing for the central claim that the queries capture culturally grounded risk signals and local norms; without these details the reported performance gaps cannot be confidently attributed to cultural misalignment rather than expert-specific perspectives.","section":"§3"},{"comment":"§4 (Evaluation and Results): The protocols for policy enforcement at inference time for dynamic models, the exact metrics used to quantify 'overestimation' relative to English-centric benchmarks, and any statistical tests for cross-variant comparisons are not specified. These omissions undermine the strength of the findings that cross-lingual transfer is insufficient and that dynamic models fail to fully localize contexts.","section":"§4"}],"minor_comments":[{"comment":"Abstract: The phrasing 'three distinct variants: static, dynamic, and multilingual' leaves unclear whether multilingual is an orthogonal dimension or overlaps with the other two; a brief clarification would improve readability.","section":"Abstract"},{"comment":"Throughout: Ensure that all tables reporting model performance include explicit definitions of the safety metrics and the exact number of queries per language or policy category.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The topic is timely and the journal scope appears appropriate, but the current lack of methodological transparency on expert validation and evaluation protocols would make it difficult for reviewers to assess reproducibility."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive and detailed feedback on our manuscript. We address each major comment below in a point-by-point manner and have revised the paper to enhance transparency where the concerns are valid.","responses":[{"response":"We agree that the original manuscript lacked sufficient detail on these aspects of the expert authoring process, which is important for supporting claims of cultural grounding. In the revised version, we have added an expanded subsection in §3 that specifies the selection criteria (domain experts recruited via African academic and professional networks with priority on representation from East, West, and Southern Africa covering major language families such as Bantu and Niger-Congo), the elicitation protocol (structured sessions with guidelines focused on local norms and field-specific harms), and the review process (iterative refinement with multiple expert reviews per query). We explicitly note the absence of formal inter-expert agreement statistics and direct community validation as limitations, while clarifying that the benchmark targets high-impact languages rather than exhaustive coverage of all 2000+ languages. These additions allow readers to better evaluate whether performance gaps reflect cultural misalignment.","revision_made":"yes","referee_comment":"[§3] §3 (Benchmark Construction): The description of the 155 domain experts' query authoring process provides no information on expert selection criteria, geographic or linguistic representation across Africa's 2000+ languages, elicitation protocol, inter-expert agreement statistics, or any external validation against community input. This is load-bearing for the central claim that the queries capture culturally grounded risk signals and local norms; without these details the reported performance gaps cannot be confidently attributed to cultural misalignment rather than expert-specific perspectives."},{"response":"We concur that the evaluation protocols and metrics required more explicit specification to strengthen the reported findings. The revised §4 now details the dynamic model policy enforcement protocol (policy text concatenated to the input prompt at inference time with reference response alignment via semantic similarity), defines the overestimation metric precisely (absolute difference in safety violation rates between English-centric benchmarks and UbuntuGuard), and includes statistical tests (Wilcoxon signed-rank tests for cross-variant comparisons with p-values added to the results tables). These clarifications directly support the conclusions on partial cross-lingual coverage and limitations of dynamic models in localizing African-language contexts.","revision_made":"yes","referee_comment":"[§4] §4 (Evaluation and Results): The protocols for policy enforcement at inference time for dynamic models, the exact metrics used to quantify 'overestimation' relative to English-centric benchmarks, and any statistical tests for cross-variant comparisons are not specified. These omissions undermine the strength of the findings that cross-lingual transfer is insufficient and that dynamic models fail to fully localize contexts."}],"tokens_in":1415,"tokens_out":573,"duration_ms":78557,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core move is to build UbuntuGuard from adversarial queries written by 155 domain experts in fields like healthcare, then turn those into context-specific policies and reference answers for evaluating guardian models in African languages. It tests 15 models across static, dynamic, and multilingual setups and reports that English-centric benchmarks overstate safety while cross-lingual transfer falls short and even dynamic models struggle to localize properly. That framing is useful because it shifts from rigid categories to runtime policies that can adapt to local norms and harm scenarios. The evaluation covers both general LLMs and dedicated guardian models, which gives a practical sense of where current systems sit. The work also flags the gap for low-resource languages explicitly rather than treating them as an afterthought. The main soft spot is the thin account of how the queries were produced and checked. The abstract notes the experts but says little about geographic or linguistic spread across Africa's many languages, elicitation methods, agreement between experts, or any external check against community views. If those queries mainly reflect the experts' professional lenses instead of wider sociocultural patterns, the performance gaps may not travel well. The full paper may fill this in, but the provided description leaves the grounding claim harder to assess. Readers working on multilingual safety or equitable model deployment will find the benchmark idea and the reported shortfalls worth seeing. The paper deserves a serious referee because it targets a real deployment gap with a concrete new resource, even if the methods section needs tightening on expert process and validation. I'd send it for review with a request for those details rather than desk-rejecting it.","headline":"UbuntuGuard brings a needed policy benchmark for African-language safety but rests on expert queries whose selection and validation get little detail.","tokens_in":2335,"tokens_out":384,"would_cite":false,"duration_ms":39520,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"UbuntuGuard AI-safety benchmark for African languages has no overlap with RS recognition-cost or forcing-chain machinery","alignment":"orthogonal","rationale":"The paper constructs a multilingual policy benchmark from expert-authored queries, derives context-specific safety rules, and evaluates static/dynamic guardian models on translation and localization tasks. Its central objects (policy-dialogue pairs, GEMBA scoring, F1 on PASS/FAIL labels) lie entirely in the domain of NLP evaluation and cultural AI safety. RS theorems (reality_from_one_distinction, J-cost uniqueness via Aczél, Alexander-duality D=3 forcing, 8-tick periodicity, phi-ladder constants) address the emergence of spacetime and physical constants from bare distinguishability; none of these structures or claims appear in or are contradicted by the UbuntuGuard construction.","tokens_in":49733,"confidence":"high","tokens_out":182,"duration_ms":17508,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"UbuntuGuard shows English AI safety benchmarks overestimate protection for African languages.","keywords":["AI safety","African languages","multilingual benchmarks","guardian models","cultural alignment","policy-based evaluation","low-resource languages"],"falsifier":"If models scored as safe on UbuntuGuard queries but still produced culturally misaligned outputs in actual African-language user tests, that result would undermine the claim that such expert-derived benchmarks are required for reliable evaluation.","tokens_in":2635,"feed_emoji":"🌍","tokens_out":610,"duration_ms":43941,"temperature":0.7,"pith_summary":"The paper creates UbuntuGuard to test AI guardian models on safety issues specific to African languages and local cultural expectations. It claims that most current benchmarks rely on English and Western categories that miss real-world harms and norms in low-resource languages. The benchmark is built from adversarial queries written by 155 domain experts in areas like healthcare, which are then turned into context-specific policies and reference answers. When 15 models are tested, the results indicate that English-centric checks give too optimistic a picture, cross-language transfer helps only partially, and even dynamic models that apply policies at runtime still fail to adapt fully to African contexts. This points to the need for benchmarks that reflect local standards so AI safety systems can work equitably across languages.","feed_headline":"New benchmark finds AI safety models overlook African cultural risks","feed_subtitle":"Expert queries in local languages show English benchmarks overestimate safety and dynamic models still miss context","key_machinery":"UbuntuGuard, the first policy-based safety benchmark for African languages, constructed from expert-authored adversarial queries that yield context-specific safety policies and reference responses for model evaluation.","core_discovery":"UbuntuGuard demonstrates that existing English-centric benchmarks overestimate real-world multilingual safety, that cross-lingual transfer supplies only partial coverage, and that dynamic guardian models, though better at using policies during inference, still cannot fully localize safety decisions to African-language cultural contexts.","pith_inferences":["Similar expert-driven benchmarks for other low-resource language groups could expose comparable shortfalls in current safety methods.","Routine involvement of community experts when building benchmarks may improve cultural fit in future AI safety work.","Models could combine dynamic policy application with targeted adaptation to specific languages for better results."],"forward_implications":["English-centric benchmarks give an inflated sense of how safe models are across languages.","Cross-lingual transfer leaves significant gaps in handling harms specific to African contexts.","Dynamic models improve on static ones by using policies at runtime but still need stronger localization.","Culturally grounded benchmarks become necessary to build guardian models that respect local expectations."],"fun_headline_variants":["English benchmarks overestimate AI safety for African languages","Guardian models miss cultural context in African language safety","Dynamic models cannot fully localize AI safety to African cultures","Cross-lingual transfer provides insufficient AI safety for African languages"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The assumption that queries written by 155 domain experts across fields including healthcare accurately reflect culturally grounded risk signals, local norms, and harm scenarios in African languages.","fun_headline_variants_meta":{"raw":{"variants":["English benchmarks overestimate AI safety for African languages","Guardian models miss cultural context in African language safety","Dynamic models cannot fully localize AI safety to African cultures","Cross-lingual transfer provides insufficient AI safety for African languages"]},"model":"grok-4.3","cost_usd":0.007214,"raw_usage":{"total_tokens":3233,"prompt_tokens":641,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":72140500,"prompt_tokens_details":{"text_tokens":641,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2533,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":641,"tokens_out":59,"duration_ms":41745,"temperature":1.0,"reasoning_tokens":2533,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-21T14:56:12.803825+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"If models scored as safe on UbuntuGuard queries but still produced culturally misaligned outputs in actual African-language user tests, that result would undermine the claim that such expert-derived benchmarks are required for reliable evaluation.","supporting_citations":[],"review_version":1}