{"id":"9f7e49a5-9395-4269-820a-5d0e3560ff1f","arxiv_id":"2411.10083","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Xmodel-1.5, a 1B multilingual LLM with a custom unigram tokenizer, outperforms PolyLM-1.7B on several Thai, Arabic, French, and Chinese benchmarks and includes a new Thai evaluation dataset.","lead":"This paper introduces Xmodel-1.5, a small 1-billion-parameter AI language model trained on 2 trillion words in multiple languages, and reports it beats a larger rival model on several non-English tests. It also releases a new Thai quiz set to help measure progress in Thai language AI.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 7 contains duplicated accuracy values across distinct tasks, undermining the multilingual comparison that supports the central claim of outperforming PolyLM-1.7B.","rationale":"The reader's weakest_assumption is that the evaluation protocol is underspecified, which is a legitimate concern about comparability across models. My concern is more concrete and more severe: Table 7 shows exact numerical duplicates across distinct tasks, which is a direct internal red flag rather than a mere lack of specification. If the duplicated values are genuine, the evaluation pipeline likely reused results or the table was populated incorrectly, so the central claim cannot be verified. If the values are typographical errors, the paper still needs a corrected table before its main comparison can be assessed. This does not impugn the authors' integrity; it simply means the current manuscript does not provide trustworthy evidence for the headline result. I acknowledge the valuable contributions of releasing the model weights and the Xdata_Thai dataset, which are separable from the contested evaluation table. The reader's conditional verdict is reasonable, but the duplicated values move me toward UNVERDICTED because the key empirical evidence is internally inconsistent and requires correction or independent reproduction before any verdict on the comparative claim can be reached.","tokens_in":13741,"tokens_out":7019,"duration_ms":64582,"concrete_test":"Re-run the evaluation for piqa_ar, piqa_fr, and xcopa_zh for each of PolyLM-1.7B, PolyLM-13B, and Xmodel-1.5-1B in the same harness, and compare the three resulting accuracies per model. If the three values are not identical for any model, Table 7 contains a copying error. If the values remain identical, inspect the evaluation scripts and per-sample predictions to verify that each dataset was loaded, formatted, and scored separately; if a shared prediction file or unreset dataset index caused the duplication, the evaluation pipeline is systematically flawed and all Table 7 results need to be regenerated.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim that Xmodel-1.5-1B outperforms PolyLM-1.7B on Thai, Arabic, French, and Chinese tasks rests on Table 7. That table contains an implausible pattern: for each model, the same accuracy value appears on three different tasks. Specifically, PolyLM-1.7B is reported as 0.5381 on piqa_ar, piqa_fr, and xcopa_zh; PolyLM-13B is reported as 0.5653 on those same three tasks; and Xmodel-1.5-1B is reported as 0.5789 on all three. These are independent datasets (Arabic PIQA, French PIQA, Chinese XCOPA), so identical four-decimal accuracies across all three are effectively impossible. This duplication indicates either a copy-paste error in the table or a systematic evaluation bug, such as accidentally reusing the same prediction file or not resetting the dataset between tasks. In either case, the numbers cannot be trusted as independent measurements. Because these rows are part of the evidence for the claimed multilingual advantage over PolyLM-1.7B, the manuscript's primary empirical claim is not reliably supported as written.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Xmodel-1.5, a 1-billion-parameter multilingual language model pretrained on 2 trillion tokens, with a custom 65,280-token unigram tokenizer. The authors describe the data mix, tokenizer design, architecture, and training details, and evaluate the model on English commonsense benchmarks, multilingual tasks (mMMLU, PIQA, XCOPA, Belebele, etc.), and a newly released Thai evaluation dataset, Xdata_Thai. The central claims are that Xmodel-1.5-1B outperforms Alibaba's PolyLM-1.7B on selected Thai, Arabic, French, and Chinese tasks, and that it achieves state-of-the-art results in Thai. The paper also reports instruction-following and chat results, and discusses qualitative feedback from a Chulalongkorn University collaboration.","tokens_in":14105,"tokens_out":4687,"duration_ms":42224,"significance":"If the reported results hold, the paper would provide a useful data point for 1B-scale multilingual models, particularly for low-resource languages like Thai. The public release of model and code, the detailed tokenizer comparison, and the construction of a human-annotated Thai evaluation dataset are concrete contributions. However, the evidence supporting the central empirical claims is currently weakened by apparent data duplications in the main multilingual results table, the absence of a uniform evaluation protocol for the benchmark comparisons, and the lack of statistical grounding for the Thai-specific superiority claim. These issues are fixable but are load-bearing for the paper's conclusions.","major_comments":[{"comment":"Table 7 reports identical accuracy values for three distinct tasks in each model row: PolyLM-1.7B has 0.5381 on piqa_ar, piqa_fr, and xcopa_zh; PolyLM-13B has 0.5653 on those same three tasks; and Xmodel-1.5-1B has 0.5789 on all three. Because these are different datasets (Arabic PIQA, French PIQA, and Chinese XCOPA), four-decimal equality across all three is effectively impossible under independent evaluation. This suggests a copy-paste error or a systematic evaluation bug (e.g., reusing the same prediction file). These rows contribute to the claimed multilingual advantage over PolyLM, so the accuracy values for these tasks must be rerun and corrected, or the affected numbers cannot be considered reliable evidence.","section":"Section 9.3, Table 7b/7c/7d"},{"comment":"The evaluation protocol for the multilingual benchmarks in Table 7 and Figure 5 is not specified. The paper gives a detailed protocol only for Xdata_Thai (Appendix 9.2.2: 3-shot continuation prompts, randomized options, first-10-token matching). For mMMLU, PIQA_AR, Belebele, XCOPA, ARC-ZH, and related tasks, the number of shots, prompt formatting (with or without chat template), answer normalization, and whether accuracy is computed by token matching or by log-likelihood are not stated. Without a common protocol across models, the comparison to PolyLM-1.7B is not reliably interpretable, because observed differences could reflect evaluation choices rather than model quality. The authors should specify the exact settings for every task used in the central comparison.","section":"Section 5.1 and Section 9.3"},{"comment":"The claim of 'state-of-the-art results in Thai' is not supported by the evidence presented. The only Thai comparisons are against PolyLM-1.7B and PolyLM-13B on Xdata_Thai, Belebele_tha, and xcopa_th. No comparison is made to other Thai-capable models (e.g., SeaLLM, Typhoon, WangchanLion, Qwen2.5) or to established Thai benchmarks. The SOTA claim in the abstract and Section 7 should either be backed by a broader comparison or removed.","section":"Abstract, Section 5.1, Section 9.3"},{"comment":"On the 350-sample Xdata_Thai dataset, the reported difference between Xmodel-1.5-1B (0.237) and PolyLM-1.7B (0.228) corresponds to roughly three to four questions, which is within the noise of a 350-item set. The paper reports no confidence intervals, error bars, or significance tests, so the conclusion that Xmodel-1.5 is 'effective' on this dataset and superior to PolyLM is not statistically established. The authors should report the number of correct answers and an appropriate uncertainty quantification or significance test.","section":"Table 6, Appendix 9.2.2"}],"minor_comments":[{"comment":"The phrase 'An 1B-scale' should be 'A 1B-scale' for grammatical correctness.","section":"Title and Abstract"},{"comment":"The comparison table includes InternLM2-1.8B and Qwen2.5-1.5B, which are 1.8B and 1.5B parameters, respectively, despite the text stating that the baselines have 'approximately 1 billion parameters.' This size mismatch should be acknowledged.","section":"Section 5.1, Table 4"},{"comment":"The compression rate in Table 2 is described as lower-is-better, but the definition of the compression rate (e.g., average tokens per word, bytes per token) is not given. Please add a one-line definition.","section":"Section 3.2, Table 2"},{"comment":"The citation for mHellaswag attributes the benchmark to Hendrycks et al. (2021), which is the MMLU paper; the HellaSwag benchmark is by Zellers et al. (2019). Please correct the reference.","section":"Section 5.1, bullet points"},{"comment":"The reference '[Johannes Welbl, 2017]' appears malformed; it should be a proper citation to Welbl et al. (2017).","section":"Appendix 9.2.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's main empirical claim is not reliably supported as written because of the duplicated values in Table 7 and the missing evaluation protocol for the multilingual benchmarks. I believe these issues are correctable: the authors could rerun the affected evaluations, provide the exact settings for every task, and soften or substantiate the Thai SOTA claim. The model and dataset releases are useful contributions, and with a careful revision the paper could become acceptable. I recommend major revision rather than rejection because the concerns are empirical and presentation-related rather than fundamental to the design or scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper reports a 1B multilingual model with a unigram tokenizer and a new Thai eval set. The recipe is fairly standard and the Thai set is only 350 samples, but the authors are honest about limitations and they release code and data, which is the right kind of behavior.\n\nThe problem is Table 7. For each model, the same accuracy appears on piqa_ar, piqa_fr, and xcopa_zh: 0.5381 for PolyLM-1.7B, 0.5653 for PolyLM-13B, and 0.5789 for Xmodel-1.5-1B. These are three independent datasets in three languages. Identical four-decimal values across all three, for all three models, is not a coincidence. It looks like a copy-paste error or a shared prediction file. Either way, the rows cannot be trusted as independent measurements. Since this table is the main evidence that Xmodel-1.5 beats PolyLM on multilingual tasks, the headline claim is not reliably supported as written.\n\nThere are smaller problems. The \"state-of-the-art in Thai\" claim rests on a single comparison with PolyLM-1.7B, not on a broad sweep of Thai benchmarks. The evaluation protocol (number of shots, prompt format, normalization, matching rule) is spelled out only for Xdata_Thai, not for the Table 7 benchmarks. And 350 samples is too small to draw strong conclusions; the reported 0.237 vs 0.228 on it is basically noise.\n\nThat said, the paper has real value. The tokenizer comparison, the training details, and the Thai dataset creation process are documented carefully. The self-citation to prior Xmodel work is not a problem. The model itself is a legitimate engineering contribution for low-resource languages, and releasing it helps.\n\nBottom line: the paper deserves a serious referee, not because it is currently correct, but because the flaw is fixable. The authors can rerun evaluations, report the exact protocol for every benchmark, and add missing baselines. If the duplication is a typo, the corrected comparison may stand. If it reflects a systematic bug, that needs to come out in review. I would not cite this version in my own work, and I would want to see the corrected table before trusting any of the multilingual numbers.","headline":"The central multilingual comparison is undermined by duplicated accuracy values in Table 7, though the model release and Thai dataset are real contributions.","tokens_in":14428,"tokens_out":2079,"would_cite":false,"duration_ms":19260,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Xmodel-1.5, a 1-billion-parameter multilingual language model, claims to outperform the larger PolyLM-1.7B on selected Thai, Arabic, French, and Chinese benchmarks and to achieve state-of-the-art results in Thai.","keywords":["multilingual language model","unigram tokenizer","low-resource languages","Thai evaluation dataset","Xdata_Thai","1B parameter model","commonsense reasoning","instruction tuning"],"falsifier":"Re-run PolyLM-1.7B on the Table 7 tasks under the exact 3-shot, token-based matching protocol used for Xmodel-1.5; if PolyLM-1.7B then matches or exceeds Xmodel-1.5's scores, the outperformance claim is falsified. Alternatively, run every publicly available Thai-capable model on Xdata_Thai under the same 3-shot prompt and show that any model scores above 0.237, falsifying the state-of-the-art-in-Thai claim.","tokens_in":13556,"feed_emoji":"🌐","tokens_out":7736,"duration_ms":59578,"temperature":0.7,"pith_summary":"Xmodel-1.5 is a 1-billion-parameter multilingual language model pretrained on 2 trillion tokens. The paper claims that this compact model, built around a custom 65,280-token unigram tokenizer, outperforms the larger PolyLM-1.7B on selected Thai, Arabic, French, and Chinese evaluation tasks, and that it achieves state-of-the-art results in Thai. It also introduces Xdata_Thai, a 350-question Thai evaluation dataset that highlights challenges like gendered particles and idioms. The authors argue that careful tokenization and targeted multilingual data can make small models competitive in low-resource languages, offering a more scalable path to multilingual AI.","feed_headline":"One-billion-parameter model tops larger PolyLM on multilingual tests","feed_subtitle":"Compact 1B model beats PolyLM-1.7B on Thai, Arabic, French, Chinese.","key_machinery":"The load-bearing mechanism is the custom unigram tokenizer. Trained with SentencePiece on a 50GB subset of the pretraining corpus (50% English, 25% Chinese, 10% industry-specific, 15% low-resource languages), it uses byte fallback for rare characters, splits numbers into digits, and keeps extra whitespace instead of removing it. The resulting 65,280-token vocabulary achieves a compression rate of 0.3800, beating LLaMA 3's 0.3823 despite having about half the vocabulary size. The tokenizer's flexibility with low-frequency tokens is what the paper argues lets a 1B model handle Thai and Arabic morphology efficiently enough to outperform a 1.7B model.","core_discovery":"The central claim is that Xmodel-1.5-1B beats PolyLM-1.7B on the paper's chosen multilingual benchmarks, including Belebele Thai, XCOPA Thai, Chinese ARC-e, and Arabic PIQA, and achieves state-of-the-art results in Thai. The paper attributes much of this to the custom unigram tokenizer: a 65,280-token SentencePiece unigram vocabulary with byte fallback, digit splitting, and no extra-whitespace removal, which reaches a compression rate of 0.3800, lower than several larger BPE vocabularies. On the released Xdata_Thai benchmark, Xmodel-1.5 scores 0.237 versus PolyLM-1.7B's 0.228 under a 3-shot setting, and the model also posts a 92.47% satisfaction rate on an e-commerce RAG evaluation after instruction tuning.","pith_inferences":["The comparison with PolyLM-1.7B could be affected by unspecified evaluation settings for Table 7, so the outperformance should be treated as provisional until the protocol is shared and reproduced.","The state-of-the-art in Thai claim rests on comparisons shown in the paper with PolyLM-1.7B and PolyLM-13B; a public leaderboard test against all Thai-capable models would be a stronger check.","The same tokenizer recipe could be tested on other low-resource languages, such as Hindi or Swahili, to see whether the compression and accuracy gains generalize.","The release of Xdata_Thai, with its focus on idioms and gendered particles, could become a targeted probe for measuring progress on culturally specific language generation."],"forward_implications":["If correct, the results show that 1-billion-parameter models can compete with 1.7B models on targeted multilingual benchmarks, making deployment cheaper and faster.","The unigram tokenizer's compression advantage suggests that tokenization design can matter as much as model scale for low-resource languages.","Xdata_Thai provides a reusable benchmark for cultural-linguistic phenomena such as gendered particles and idioms, which standard multilingual benchmarks do not cover.","The strong e-commerce RAG performance (92.47% satisfaction) indicates the instruction-tuned model is usable for real commercial multilingual customer service."],"supporting_citations":[{"why":"Provides PolyLM, the 1.7B and 13B multilingual baseline that Xmodel-1.5 is directly compared against on Thai, Arabic, French, and Chinese tasks.","marker":"[Wei et al., 2023]"},{"why":"Supplies the subword regularization / unigram language model method that the custom tokenizer is based on.","marker":"[Kudo, 2018a]"},{"why":"Supplies SentencePiece, the implementation used to train the unigram tokenizer.","marker":"[Kudo and Richardson, 2018]"},{"why":"Provides CulturaX, the main source of multilingual pretraining data for low-resource languages.","marker":"[Nguyen et al., 2023]"},{"why":"Supplies XCOPA, the multilingual commonsense reasoning benchmark used for Thai and Chinese evaluation.","marker":"[Ponti et al., 2020]"},{"why":"Supplies Belebele, the Thai reading comprehension benchmark that underpins the Thai performance claim.","marker":"[Bandarkar et al., 2023]"},{"why":"Supplies PIQA_AR, the Arabic physical commonsense benchmark used in the Arabic evaluation.","marker":"[Almazrouei et al., 2023]"},{"why":"Provides the Xmodel-LM architecture and data that Xmodel-1.5 builds upon.","marker":"[Wang et al., 2024a]"}],"fun_headline_variants":["1B model outperforms PolyLM-1.7B on multilingual tests","Custom unigram tokenizer gives 1B LLM edge over PolyLM","Xmodel-1.5: small model, big multilingual wins","A 1B LLM beats a 1.7B rival across five languages","State-of-the-art Thai from a 1B model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the multilingual evaluations used the same prompt format, number of shots, normalization, and token-based matching for Xmodel-1.5 and PolyLM-1.7B; the paper spells out those settings only for its own Xdata_Thai benchmark, not for the Table 7 results.","fun_headline_variants_meta":{"raw":{"variants":["1B model outperforms PolyLM-1.7B on multilingual tests","Custom unigram tokenizer gives 1B LLM edge over PolyLM","Xmodel-1.5: small model, big multilingual wins","A 1B LLM beats a 1.7B rival across five languages","State-of-the-art Thai from a 1B model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000304,"raw_usage":{"total_tokens":1740,"prompt_tokens":932,"completion_tokens":808,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":711}},"tokens_in":548,"tokens_out":808,"duration_ms":7237,"temperature":1.0,"reasoning_tokens":711,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T19:58:36.600798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run PolyLM-1.7B on the Table 7 tasks under the exact 3-shot, token-based matching protocol used for Xmodel-1.5; if PolyLM-1.7B then matches or exceeds Xmodel-1.5's scores, the outperformance claim is falsified. Alternatively, run every publicly available Thai-capable model on Xdata_Thai under the same 3-shot prompt and show that any model scores above 0.237, falsifying the state-of-the-art-in-Thai claim.","supporting_citations":[{"cited_title":"Rossi, and Thien Huu Nguyen","cited_arxiv_id":null,"evidence_quote":"Provides CulturaX, the main source of multilingual pretraining data for low-resource languages."},{"cited_title":"The belebele benchmark: a parallel reading comprehension dataset in 122 language variants, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies Belebele, the Thai reading comprehension benchmark that underpins the Thai performance claim."}],"review_version":1}