{"id":"71373756-616c-4008-ba4b-3f4d113cddde","arxiv_id":"2504.15527","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Compass-v2 is an efficient MoE language model for Southeast Asian languages and e-commerce that claims top-tier performance among models of its size using a unified fast/deep reasoning design.","lead":"This report introduces Compass-v2, a 30-billion-parameter mixture-of-experts language model that activates only 5 billion parameters per token and targets Southeast Asian languages and e-commerce. It claims state-of-the-art multilingual and e-commerce performance for its size while cutting inference cost, and it pairs fast and deep reasoning in one model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified","rationale":"The reader's weakest assumption---that the in-house evaluation is fair and valid---is exactly where the paper's central claim is least secure. The paper states in Section 7.5 that in-house free-form benchmarks are scored by GPT-4, with no released prompts, no contamination analysis against the SFT/DPO data in Sections 3-4, and no error bars in any table. The strongest claim ('state-of-the-art SEA multilingual and e-commerce performance among sub-30B models') rests substantially on Table 9, which shows Compass-v2 averaging 76.64 versus Qwen2.5-32B at 77.00 and GPT-4o at 77.68---a claim of 'comparable' rather than 'superior.' The lack of public artifacts is a genuine verification gap, but it is not an internal inconsistency or a demonstrated error. The open-source benchmark tables are internally consistent and plausible, and the paper's own numbers do not contradict each other. Therefore the reader's CONDITIONAL verdict is appropriate and no adjustment is needed. I agree with the reader's identification of the weakest assumption, and my concern is the same one, framed as a verification gap rather than a demonstrated flaw.","tokens_in":30626,"tokens_out":1335,"duration_ms":11299,"concrete_test":"Request or independently reproduce the in-house evaluation: obtain the exact prompts from Section 7.5 (or the released evaluation harness), run Compass-v2 alongside Qwen2.5-14B/32B and GPT-4o with identical sampling settings, and compute per-language scores with bootstrap confidence intervals. If the released scores are not reproducible or the margin over Qwen2.5-32B closes within noise, the 'comparable to GPT-4o' claim weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim---state-of-the-art SEA multilingual and e-commerce performance among sub-30B models at 5B active parameters---is a commercial self-report with no public model, data, or evaluation harness. That absence makes independent verification impossible, but it is a limitation of the artifact, not an internal inconsistency. The open-source benchmark tables (Tables 6-8) show plausible, internally consistent numbers, and the in-house evaluation (Table 9) is explicitly framed as Shopee-internal with GPT-4 as judge. The reader's concern about contamination and judge bias is real but is a verification gap: no evidence in the paper contradicts the claimed results, and no evidence confirms them. The weaker spot is the in-house evaluation's reliance on GPT-4 as judge over tasks built from the same business domain used for SFT and DPO, with no released prompts or error bars (Section 7.5). This is the most load-bearing because the headline 'state-of-the-art' claim depends heavily on Table 9, but the lack of external checkpoints means the claim rests on the authors' own evaluation protocol. I do not find an internal contradiction that would justify rejection; the appropriate disposition is a conditional acceptance pending release of evaluation artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This technical report describes Compass-v2, a fine-grained Mixture-of-Experts (MoE) language model with 30B total and 5B active parameters, designed for Southeast Asian (SEA) languages and e-commerce applications. The authors present a three-stage pretraining pipeline on 12T tokens (generic pretraining, high-quality annealing, and long-context extension), a two-stage SFT on about 6M instructions, and an online-DPO alignment stage. They also introduce a hybrid reasoning design that supports both fast-thinking and deep-thinking responses within one model. Evaluation is reported on open multilingual, e-commerce, and general English benchmarks, and on in-house Shopee datasets scored by GPT-4. The abstract claims state-of-the-art SEA multilingual and e-commerce performance among sub-30B models with significantly lower inference cost.","tokens_in":30837,"tokens_out":10116,"duration_ms":79015,"significance":"If the reported results are accurate, Compass-v2 is a practically relevant contribution: a compact MoE trained from scratch for SEA languages and e-commerce, with a tokenizer that achieves strong compression on SEA languages (Figure 6) and a hybrid reasoning mode that reportedly preserves general performance within 1% (Section 5.3). The open-benchmark tables (Tables 6-8) provide some external grounding and show a broadly competitive model, and the data-quality comparison (Table 1) is a useful reference point. However, the absence of a public model, data, or evaluation harness, together with the reliance on a private GPT-4-scored evaluation for the strongest claims, limits the immediate impact of the work until those artifacts are released.","major_comments":[{"comment":"The abstract and Section 1 claim 'state-of-the-art SEA multilingual and e-commerce performance among sub-30B models.' This is not fully supported by the paper's own open-benchmark tables. In Table 8, Compass-v2 scores 0.7801 on MMLU, below Qwen2.5-14b's 0.7991; 0.4091 on GPQA, below Qwen2.5-14b's 0.4394; 0.5913 on ARC, below Qwen2.5-14b's 0.6220; and 0.8150 on GSM8K, below Qwen2.5-7b's 0.8347. Since Qwen2.5-14b and Qwen2.5-7b are both sub-30B models, the claim either needs a clearly defined aggregation rule under which Compass-v2 is indeed the best, or a more modest phrasing such as 'competitive with' or 'among the best.'","section":"Abstract, §1, §7.4, Table 8"},{"comment":"The in-house evaluation is load-bearing for the claims of state-of-the-art SEA multilingual and e-commerce performance and for the statement that Compass-v2 is 'comparable' to GPT-4o and Qwen2.5-72b. However, this evaluation is not externally verifiable: the prompts are not released, no contamination analysis is performed against the SFT and DPO data described in Sections 3 and 4, and the judge is GPT-4, which is also one of the baselines in Table 9. Moreover, no error bars or statistical tests are reported for any of the tables. To support the central claims, the authors should release the evaluation prompts (or a substantial sample), provide a contamination check, describe the judge protocol in detail, and report variance across repeated evaluations.","section":"§7.5, Table 9"},{"comment":"The abstract claims that the authors 'pioneered a hybrid reasoning model that supports both fast thinking and deep thinking within a unified framework.' However, Section 5.1 explicitly cites Anthropic (2025) as a prior hybrid reasoning model of this kind. The word 'pioneered' is therefore inaccurate and should be replaced with a more measured term such as 'introduced' or 'presented.' The novelty of the hybrid reasoning design relative to Anthropic (2025) should also be stated explicitly.","section":"Abstract, §5.1"},{"comment":"The claim that the Compass-v2 (EN) dataset 'surpasses the best open-source datasets in the industry' rests on a single evaluation run with no reported variance. The difference between Compass-v2 (EN) (0.6397) and Combined (EN) (0.6366) is only 0.0031 in average score, which is within typical run-to-run noise for a 210B-token pretraining comparison. The paper should report multiple seeds, confidence intervals, or a significance test before making this strong claim.","section":"§2.1.2, Table 1"}],"minor_comments":[{"comment":"There are several typos: 'Comapss-v2' (Section 1), 'Quantizaiton' (Section 1), 'Qwem2.5-72b' (Section 1), 'wtih' (Section 7.4), 'pionneering' (Section 8.1), 'instend' (Section 4.1), and 'reuslt' (Section 3.1.1). Please copyedit the manuscript.","section":"Throughout"},{"comment":"The auxiliary loss formula for load balancing is typeset as Laux = N · Σ_i (p_i/B · c_i/B · K), which is dimensionally unclear. Please provide the correct mathematical expression or explicitly define p_i, c_i, B, and K so that the formula is unambiguous.","section":"§2.2"},{"comment":"The acronym 'ABF' is used without expansion. Please define it on first use (e.g., 'adjusted base frequency,' citing Xiong et al., 2023).","section":"§2.3.3"},{"comment":"The evaluation settings for baselines are incomplete; for example, Sailor2 and Moonlight are given as 'huggingface opensource generation config.' For reproducibility, consistent generation parameters should be listed for all baselines.","section":"§7.1"},{"comment":"The disclaimer in Section 10 is a legal boilerplate and is not a substitute for a scientific limitations statement. Please add a short limitations paragraph addressing the internal-evaluation constraints and the absence of public artifacts.","section":"§10"}],"recommendation":"major_revision","confidential_remarks":"The paper is a technical report rather than a full research paper. Its most striking claims (e.g., beating GPT-4o on Portuguese in Table 9) rest entirely on an internal, GPT-4-scored evaluation with no released prompts, no contamination check, and no error bars. The open-benchmark tables are more reassuring but do not by themselves support the strongest abstract claims. I recommend that the editor weigh whether the journal's reproducibility standards are compatible with a submission whose central evidence is a private Shopee-internal evaluation, and whether the authors' own disclaimer (Section 10) is consistent with the paper's stated conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick read on arXiv:2504.15527. It is an industry tech report for Compass-v2, a 30B-total/5B-active MoE trained from scratch on 12T tokens for SEA languages and e-commerce. The genuinely new parts are the assembled system, the SEA-adapted tokenizer, the in-house corpus, and the single-model hybrid fast/deep reasoning setup. The architecture pieces are borrowed (DeepSeekMoE shared-plus-fine-grained experts, Claude 3.7-style hybrid reasoning), but the engineering is substantial and the open-benchmark tables are internally consistent. I found the data-curation ablation in Table 1 (more dedup hurts) worth reading, and the long-context needle and quantization measurements are useful.\n\nThe soft spots are real but not fatal. No model, data, or evaluation harness is released, so everything rests on the authors' word. The in-house evaluation (Table 9) uses GPT-4 as judge on Shopee-domain tasks built from the same business area as the SFT/DPO data, with no contamination analysis and no error bars. The 'sub-30B' framing is confusing for a model with 30B total parameters; the comparisons include Qwen2.5-72b and GPT-4o, so the headline must be read as active-parameter SOTA, not total-parameter SOTA. There is no controlled ablation isolating the hybrid-reasoning contribution; the downsampled Long-CoT compromise is described but not quantified in a table. And Section 10's disclaimer explicitly disclaims the correctness of the data and the adequacy of the methodology, which is unusual for a report meant to be read as evidence.\n\nI would not call this a useless paper. It is a credible window into how Shopee builds a domain-specific MoE, and the measured tokenizer compression rates are a contribution. But as a scientific artifact it is a commercial self-report.\n\nBottom line: it deserves a serious referee, but with conditions. Require release of the model or at least the evaluation harness, contamination analysis, uncertainty reporting, and a fix to the 'sub-30B' framing. Until then, the SOTA claim is a claim, not a result.","headline":"Real MoE engineering, no artifacts; the SOTA claim is a claim, not a result.","tokens_in":31382,"tokens_out":3469,"would_cite":false,"duration_ms":30961,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 5B-active-parameter model claims state-of-the-art Southeast Asian language and e-commerce performance among sub-30B models.","keywords":["Mixture-of-Experts","Southeast Asian languages","e-commerce","hybrid reasoning","low-resource NLP","tokenization","direct preference optimization","model quantization"],"falsifier":"Release the in-house evaluation prompts and answers, then run a near-duplicate search against the 12T-token pretraining corpus and the SFT and DPO data; finding substantial overlap would show the reported in-house scores measure memorization rather than capability. Alternatively, a rerun of the open-source benchmarks (Shopping MMLU, ECInstruct, and the general set) under identical generation settings without the custom tokenizer would test whether the advantage survives fair comparison.","tokens_in":30386,"feed_emoji":"🧭","tokens_out":6775,"duration_ms":50780,"temperature":0.7,"pith_summary":"Compass-v2 is a Mixture-of-Experts language model with 30B total parameters but only 5B active per token, trained from scratch specifically for Southeast Asian languages and e-commerce. The paper argues that careful data curation—a 12T-token corpus emphasizing Indonesian, Thai, Vietnamese, Malay, Tagalog, and Portuguese plus hundreds of billions of e-commerce tokens—and a tokenizer optimized for SEA compression let a small model beat larger dense rivals on SEA and shopping tasks. It also introduces a hybrid reasoning design that gives both fast and deep thinking from one model, instead of deploying two separate systems. If the results hold, sub-30B models trained from scratch on region-specific data can outperform general-purpose models at a fraction of the inference cost.","feed_headline":"SEA-focused MoE claims top scores with just 5B active params","feed_subtitle":"Trained from scratch on 12T tokens, it unifies fast and deep thinking to cut e-commerce inference cost.","key_machinery":"The engine of the design is a fine-grained Mixture-of-Experts (MoE) transformer in which every token routes through 2 shared experts and 4 specialized experts drawn from 48, giving 30B total parameters with only 5B active per token; a load-balancing auxiliary loss and a router Z-loss keep expert usage even during training. Supporting machinery includes a three-way sub-tokenizer merged into a 180k-vocabulary BPE tokenizer with the best reported compression on SEA languages, a three-stage pretraining schedule (8T tokens, 4T high-quality annealing, then extension from 4k to 32k context), and post-training with two SFT stages plus Online-DPO. The hybrid reasoning capability comes from long chain-of-thought data with thinking tags added in the second SFT stage, selectively downsampled to avoid degrading general skills.","core_discovery":"On its own terms, the paper reports that Compass-v2 achieves the highest average scores among comparable sub-30B models on open-source e-commerce benchmarks (55.74% average, beating the runner-up by a relative 10.6%) and on general English benchmarks (72.21% average), while activating only 5B parameters per token. It also reports in-house multilingual scores averaging 76.64, within about two points of GPT-4o and Qwen2.5-72b despite its smaller scale. The model unifies fast and deep reasoning through a single checkpoint: a long chain-of-thought prompt template triggers step-by-step thinking, and the authors report the hybrid training preserves general performance within a 1% drop. These claims rest on internal evaluation sets built from real Shopee scenarios, scored with GPT-4.","pith_inferences":["If the tokenizer compression advantage generalizes, the same sub-tokenizer merging recipe could be applied to other low-resource language families (for example, South Asian or African languages) with similar efficiency gains.","The reported results suggest that for domain-specific low-resource deployments, investing in cleaned native corpora and a bespoke tokenizer may yield more value than starting from an English-centric checkpoint, though matched compute budgets would be needed to confirm this.","A direct test of the hybrid reasoning claim would be to measure whether the same checkpoint, with the general template active, matches a dedicated fast model on latency-sensitive e-commerce queries while the LongCoT template matches a dedicated reasoning model on math benchmarks.","Contamination is the key threat: because the in-house evaluation uses unreleased Shopee-derived prompts, publishing them and scanning the 12T-token corpus for near-duplicates would let the community verify the claim."],"forward_implications":["A region-focused MoE trained from scratch can match or beat general-purpose dense models several times its active size on low-resource language tasks.","One checkpoint can serve both lightweight chat and long step-by-step reasoning, removing the operational cost of maintaining two models.","4-bit AWQ quantization preserves benchmark accuracy while delivering a 1.58x throughput gain over FP16 at high concurrency, making such models deployable on modest GPU fleets.","A tokenizer built by merging language-group sub-tokenizers improves compression for SEA languages, directly cutting per-token inference cost."],"supporting_citations":[{"why":"Supplies the sparsely-gated MoE routing mechanism that Compass-v2's expert assignment builds on.","marker":"(Shazeer et al., 2017)"},{"why":"Provides the fine-grained expert and shared-expert design that lets the model activate 5B of 30B parameters.","marker":"(Dai et al., 2024)"},{"why":"DCLM is used as the open-source dataset baseline in the 210B-token ablation showing Compass-v2's English data quality.","marker":"(Li et al., 2024b)"},{"why":"FineWeb-Edu is part of the Combined (EN) mixture used in the same data-quality comparison.","marker":"(Penedo et al., 2024)"},{"why":"Nemotron-CC is the other component of the Combined (EN) baseline in the data-quality comparison.","marker":"(Su et al., 2024a)"},{"why":"The first-generation Compass model whose curation pipeline Compass-v2 extends and compares against.","marker":"(Maria, 2024)"},{"why":"Shopping MMLU is the open-source e-commerce benchmark central to the claimed state-of-the-art result.","marker":"(Jin et al., 2024)"},{"why":"DeepSeek-R1 is used to generate the long chain-of-thought data for hybrid reasoning training.","marker":"(Guo et al., 2025)"},{"why":"Provides the DPO objective used in the Online-DPO alignment stage.","marker":"(Rafailov et al., 2024)"}],"fun_headline_variants":["SEA's Compass-v2 beats sub-30B rivals with 5B active params","MoE for SEA and e-commerce: 30B total, 5B active, tops benchmarks","Compass-v2: unified fast/deep thinking in a 5B-active MoE","SEA e-commerce MoE: 5B active params, top scores under 30B","Compass-v2 claims SOTA SEA and e-commerce with 5B active"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim of state-of-the-art performance depends on internal evaluation datasets that are not publicly released and have not been checked for overlap with the training data; if those datasets leak into pretraining or fine-tuning, or if the GPT-4 judge is biased, the reported advantage may not be real.","fun_headline_variants_meta":{"raw":{"variants":["SEA's Compass-v2 beats sub-30B rivals with 5B active params","MoE for SEA and e-commerce: 30B total, 5B active, tops benchmarks","Compass-v2: unified fast/deep thinking in a 5B-active MoE","SEA e-commerce MoE: 5B active params, top scores under 30B","Compass-v2 claims SOTA SEA and e-commerce with 5B active"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2390,"prompt_tokens":922,"completion_tokens":1468,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":1350}},"tokens_in":538,"tokens_out":1468,"duration_ms":9406,"temperature":1.0,"reasoning_tokens":1350,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:24:05.352559+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Release the in-house evaluation prompts and answers, then run a near-duplicate search against the 12T-token pretraining corpus and the SFT and DPO data; finding substantial overlap would show the reported in-house scores measure memorization rather than capability. Alternatively, a rerun of the open-source benchmarks (Shopping MMLU, ECInstruct, and the general set) under identical generation settings without the custom tokenizer would test whether the advantage survives fair comparison.","supporting_citations":[{"cited_title":"The fineweb datasets: Decanting the web for the finest text data at scale","cited_arxiv_id":null,"evidence_quote":"FineWeb-Edu is part of the Combined (EN) mixture used in the same data-quality comparison."}],"review_version":1}