{"id":"dda9691e-7fba-42f4-ae2c-e45e37ca4a5f","arxiv_id":"2607.04581","paper_version":2,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"A native 22-task Brazilian-Portuguese embedding benchmark cleanly tiers 93 models, places an open model in the unresolved top tier, and finds only moderate rank correlation (ρ=0.75) with the multilingual MTEB board.","lead":"MTEB-BR is a 22-task native Brazilian-Portuguese text-embedding benchmark that excludes machine translations by design and evaluates 93 models. It shows open-weight models reach the commercial frontier and that multilingual leaderboard ranks only moderately predict Portuguese ranks.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The strongest claim is an empirical measurement (ρ=0.75 + open model in frontier tier) directly supported by the public score matrix, bootstrap tests, and Appendix F pairwise p-values showing the top six unresolved. The reader's weakest assumption correctly flags the most important scope caveats, yet the paper already quantifies them and shows headline rankings are robust (τ=0.91 Borda, τ=0.97 after redundancy drop, discrimination screen leaves Spearman 0.99). Because the limitations are candidly bounded and do not reverse the measured moderate transfer or open-model parity, they do not justify moving the verdict from ACCEPT. The concrete test above is a useful sensitivity check rather than a required fix. Agreement with the reader is therefore full; no adjustment is warranted.","tokens_in":31118,"tokens_out":517,"duration_ms":5153,"concrete_test":"Recompute the 55-model Spearman ρ after (i) dropping the six legal/regulatory formulations (JurisTCU×3, BRTaxQAR, PortuLexRRIP, FaqBacen) and (ii) re-encoding BR-TaxQA-R at 2048 tokens for long-context models; if ρ rises above 0.90 or the top-six open/closed membership changes, the native-value claim would need re-scoping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claims (moderate cross-leaderboard Spearman ρ=0.75 over 55 shared models with large rank inversions such as Llama-Embed-Nemotron-8B 3rd→49th; open-weight Qwen3-Embedding-8B inside the unresolved top-six frontier tier) rest on a released 93-model score matrix, 10k-task bootstrap CIs, paired significance, and explicit native-source construction. The reader's weakest assumption (register skew, 512-token cap, survivor bias from the discrimination screen, web-text contamination) is already enumerated by the authors in §III, §VIII, and §X, with quantitative mitigations (mean-centered inter-task cosine ≈−0.04, Kendall τ=0.97 after dropping shared-corpus reformulations, Serafim IR-tuned gap ≈0). These are real limitations of scope, not internal contradictions that overturn the measured moderate correlation or the open-model parity result. No load-bearing flaw in the argument is apparent.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark of 22 tasks across seven MTEB categories, constructed under a strict native-source filter that excludes machine-translated corpora (notably mMARCO-PT). It evaluates 93 models (73 open-weight, 20 closed APIs; 23M–27B parameters) and reports a statistical layer: per-task bootstrap CIs, paired-bootstrap significance and tiers, task- and instance-level IRT discrimination, Borda robustness, and a cross-leaderboard Spearman correlation. Headline results are that the suite separates roughly a dozen model tiers while leaving the top six unresolved; an Apache-2.0 open model (Qwen3-Embedding-8B) sits in that frontier tier; and multilingual MTEB rank predicts Portuguese rank only moderately (ρ=0.75 over 55 shared models, with large inversions such as Llama-Embed-Nemotron-8B 3rd→49th). Code, tasks, results, and a public leaderboard are released.","tokens_in":31451,"tokens_out":1285,"duration_ms":20067,"significance":"The work fills a clear gap: Portuguese embedding evaluation has relied on translated or thinly covered multilingual suites. The native construction, large and mixed open/closed model panel, and especially the transferable statistical layer (bootstrap tiers, IRT discrimination, Borda check) are genuine contributions that other language-specific MTEB extensions can reuse. The open-model parity result and the moderate cross-leaderboard correlation are practically actionable and are backed by a released 93×22 score matrix with explicit uncertainty. Full open release of code (Apache-2.0), results (CC-BY 4.0), and an interactive leaderboard further strengthens the contribution.","major_comments":[{"comment":"§X and Appendix A: the uniform 512-token cap is acknowledged, and a single-model check on BR-TaxQA-R shows a large nDCG@10 lift when the cap is raised. Because retrieval is the dominant separator (§VIII, Table IV) and a frontier endpoint (voyage-context-4) is a long-context model, the cap conditions both the retrieval ranking and the cost–Pareto of §IX. A short multi-model ablation (or at least reporting scores at 512 vs. native context for the long-context APIs and open models that support it) would make the frontier claim more robust rather than leaving the penalty as a single-model bound.","section":"§X Limitations; §IX; Appendix A"},{"comment":"§III exclusions and §VIII: the discrimination screen that drops low-separating candidates is correctly flagged as biasing reported a_t upward, and the authors note Spearman 0.99 before/after the final cut. For a benchmark paper whose design claim rests on native + discriminating tasks, a brief appendix table of the dropped candidates (name, category, reason, and if available a_t or score variance) would let readers judge survivor bias quantitatively rather than only via the authors’ qualitative screen description.","section":"§III; §VIII"},{"comment":"§VII: the cross-leaderboard ρ=0.75 is a central claim and is well supported by the rank-inversion example and category-wise breakdowns. The live HuggingFace board is a drifting, differently composed reference; the paper already notes this. A one-sentence sensitivity check (e.g., correlation restricted to models with stable identifiers, or against the published MMTEB headline ranking already mentioned) would further insulate the “native boards measure something different” conclusion from snapshot dependence.","section":"§VII"}],"minor_comments":[{"comment":"Table III caption and §VI: score CIs are correctly described as marginal, not paired; consider adding a one-line pointer in the main text to Appendix F so readers do not misread overlapping CIs as non-significance.","section":"Table III; §VI"},{"comment":"Figure 1 and the diversity diagnostic: mean-centered cosine is the right choice; stating the three-model panel and mean pairwise Spearman ρ=0.82 in the figure caption (not only the body) would make the plot self-contained.","section":"Figure 1; §III"},{"comment":"Table II is dense (two-panel metadata for 93 models). A machine-readable companion (already released) is fine; in the PDF, a short family-level summary table or sorting by family would improve scannability.","section":"Table II"},{"comment":"§VIII-A instance-level IRT: the efficiency and misfit results are carefully scoped as a case study on four dichotomous tasks. Ensure the abstract and conclusion do not imply suite-wide adaptive evaluation without that qualifier.","section":"§VIII-A; Abstract; §XI"},{"comment":"Minor typography: “S ˜ao Paulo”, “Avaliac ¸˜ao”, and similar accent/spacing artifacts appear in affiliations and references; a pass with a Portuguese-aware PDF toolchain would clean them.","section":"Title page; References"},{"comment":"Appendix D (mMARCO train–test overlap) is valuable supporting evidence for the native-source filter; a single sentence in §III pointing to the cohort rank-shift figure would help readers who skip appendices.","section":"§III; Appendix D"}],"recommendation":"minor_revision","confidential_remarks":"Strong empirical benchmark paper with unusually careful statistics and full release. Fit for a methods/resources track or main conference/journal in CL/IR. The three major points are clarification/robustness requests, not soundness failures; I would not block acceptance over them. No novelty or citation concerns."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the missing native MTEB for Brazilian Portuguese, done carefully. What is actually new is not “another language extension” — those already exist for Chinese, French, Scandinavian, Polish, etc. — but the strict native-only filter (mMARCO-PT and mkqa-PT out by construction), a 93-model panel that includes 20 closed APIs, and a statistical layer that most language-specific boards still skip: 10k-task bootstrap CIs, paired significance with tiers, Borda robustness (τ=0.91), and both task- and instance-level IRT discrimination. The three headlines are supported by the score matrix: roughly a dozen resolvable tiers with an unresolved top six, Qwen3-Embedding-8B inside that frontier, and only moderate rank agreement with the live multilingual board (ρ=0.75 over 55 shared models, with the concrete 3rd→49th inversion on Llama-Embed-Nemotron).\n\nThe paper does the unglamorous work well. Full code, Parquet results, revision SHAs, random baseline, inter-task content and ranking diagnostics, and an appendix that shows translated mMARCO inflates scores for MS-MARCO-trained models. Limitations are listed without spin: 512-token cap, short documents, legal/institutional register skew, small retrieval pools, closed-API drift, survivor bias from the discrimination screen, possible web-text contamination. Those are real scope limits; they do not invent an order among the top six or manufacture the moderate cross-board correlation. The IRT fits are correctly labeled descriptive least-squares, not psychometric claims.\n\nSoft spots are proportionate. The discrimination screen biases a_t upward (authors say so). The Portuguese-vs-multilingual retrieval gap is underpowered once restricted to IR-tuned Serafim. Register and length bias mean the suite speaks more to legal/FAQ/search than dialogue or long-document RAG. None of that overturns the central empirical results.\n\nThis is for people who ship Portuguese retrieval or choose embedding APIs, and for anyone building the next language-specific MTEB who wants a reusable uncertainty template. I would send it to peer review; the artifacts and the measurement claims are strong enough to deserve referee time. Engage with it if you work on multilingual embeddings or Portuguese IR.","headline":"Solid native Portuguese embedding benchmark with a real statistical layer; the moderate multilingual-proxy result and open-model parity claim hold up on the released matrix.","tokens_in":32027,"tokens_out":573,"would_cite":true,"duration_ms":6394,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A native Brazilian-Portuguese embedding benchmark shows multilingual leaderboards only moderately predict Portuguese rank, and an open self-hostable model reaches the unresolved top tier.","keywords":["text embeddings","Brazilian Portuguese","MTEB-BR","native benchmark","retrieval","Item Response Theory","open-weight models","cross-leaderboard correlation"],"falsifier":"If, on a larger set of native Brazilian-Portuguese retrieval and dialogue tasks with long documents and no legal over-weighting, the same 55 models’ multilingual ranks correlated near 0.95 with the new mean and no open model reached the commercial frontier, the claim that native evaluation measures something distinct and that open models already match would fail.","tokens_in":32016,"feed_emoji":"🇧🇷","tokens_out":1017,"duration_ms":11789,"temperature":0.7,"pith_summary":"Until now, people choosing sentence-embedding models for Brazilian Portuguese had to rely on translated English sets or thin multilingual coverage. This paper builds MTEB-BR: 22 tasks drawn only from data created or found in Portuguese, spanning classification, similarity, clustering, retrieval, and reranking, and evaluates 93 models from tiny open encoders to large commercial APIs. With bootstrap intervals, paired significance, and an Item Response Theory view of how sharply each task separates models, the suite cleanly orders most of the field into about a dozen tiers while leaving the top six statistically tied. An openly licensed, self-hostable model sits in that leading cluster, so frontier Portuguese quality does not require a paid API. A model’s rank on the global multilingual board predicts its Portuguese rank only moderately (Spearman ρ = 0.75), and one model that ranks third there falls near the bottom here, so a native benchmark measures something the multilingual boards do not.","feed_headline":"Open model hits top Portuguese embedding tier; global ranks mislead","feed_subtitle":"Native 22-task board only moderately tracks multilingual leaderboards, so Portuguese needs its own evidence.","key_machinery":"MTEB-BR’s 22-task mean over native-only Portuguese tasks, backed by a statistical layer of per-task bootstrap confidence intervals, paired-bootstrap significance (tiers and probability-of-best), Item Response Theory task- and instance-level discrimination, and Borda robustness, which together turn the leaderboard into an uncertainty-aware measurement rather than a bare ranking.","core_discovery":"On 22 native Brazilian-Portuguese embedding tasks, multilingual leaderboard rank predicts Portuguese rank only moderately (Spearman ρ = 0.75 over 55 shared models; one model is 3rd there and 49th here), while an open self-hostable model reaches the statistically unresolved top tier of six, so strong Portuguese embedding quality does not require a commercial API and native evaluation is not redundant with global boards.","pith_inferences":["If other languages show similar moderate cross-leaderboard correlations, native MTEB-style boards may be necessary wherever deployment text diverges from English-mined and translated retrieval.","The training-objective account of the Portuguese-encoder retrieval gap suggests fine-tuning strong multilingual or decoder backbones on native Portuguese retrieval data is a more direct path than more language-specific pretraining alone.","Instance-level IRT recovery of rankings from the most informative examples points toward cheaper ongoing leaderboard maintenance as new models appear.","Silent closed-API drift and short-document caps mean published commercial ranks will need periodic re-runs against the same native suite before production lock-in."],"forward_implications":["Practitioners can pick Portuguese embedding models from native evidence and treat multilingual rank as a first filter only, not a proxy.","Within the top six, choice should turn on cost, license, latency, and context length rather than score, since the benchmark does not resolve their order.","Open Apache-2.0 models at the frontier remove per-token cost for self-hosting without sacrificing the quality tier measured here.","Retrieval tasks carry the sharpest separation signal; budgets for harder Portuguese evaluation are better spent there than on weak clustering or low-signal classification probes.","The same statistical layer (bootstrap tiers, paired significance, IRT discrimination) can be reused for other language-specific embedding suites."],"fun_headline_variants":["Open model reaches top Portuguese embed tier; global ranks mislead","Native BR board separates models; open one hits unresolved top six","Multilingual rank predicts Portuguese only moderately (ρ=0.75)","Self-hostable open model joins leading Portuguese embedding tier","Portuguese embed quality needs native board, not just global ranks"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The filter that keeps only native Portuguese sources and then drops low-discrimination candidates yields a 22-task mean that truly reflects Portuguese embedding quality rather than legal-register skew, short-document bias under a 512-token cap, or memorized web text.","fun_headline_variants_meta":{"raw":{"variants":["Open model reaches top Portuguese embed tier; global ranks mislead","Native BR board separates models; open one hits unresolved top six","Multilingual rank predicts Portuguese only moderately (ρ=0.75)","Self-hostable open model joins leading Portuguese embedding tier","Portuguese embed quality needs native board, not just global ranks"]},"model":"grok-4.5","effort":"low","cost_usd":0.005378,"raw_usage":{"total_tokens":1491,"prompt_tokens":844,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":53780000,"prompt_tokens_details":{"text_tokens":844,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":578,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":844,"tokens_out":69,"duration_ms":4137,"temperature":1.0,"reasoning_tokens":578,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T16:54:30.249232+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If, on a larger set of native Brazilian-Portuguese retrieval and dialogue tasks with long documents and no legal over-weighting, the same 55 models’ multilingual ranks correlated near 0.95 with the new mean and no open model reached the commercial frontier, the claim that native evaluation measures something distinct and that open models already match would fail.","supporting_citations":[],"review_version":1}