Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

A native Brazilian-Portuguese embedding benchmark shows multilingual leaderboards only moderately predict Portuguese rank, and an open self-hostable model reaches the unresolved top tier.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-11 16:54 UTC pith:ZDXMQY6R

load-bearing objection Solid native Portuguese embedding benchmark with a real statistical layer; the moderate multilingual-proxy result and open-model parity claim hold up on the released matrix. the 3 major comments →

arxiv 2607.04581 v2 pith:ZDXMQY6R submitted 2026-07-06 cs.CL cs.IRcs.LG

MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese

classification cs.CL cs.IRcs.LG
keywords text embeddingsBrazilian PortugueseMTEB-BRnative benchmarkretrievalItem Response Theoryopen-weight modelscross-leaderboard correlation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Until now, people choosing sentence-embedding models for Brazilian Portuguese had to rely on translated English sets or thin multilingual coverage. This paper builds MTEB-BR: 22 tasks drawn only from data created or found in Portuguese, spanning classification, similarity, clustering, retrieval, and reranking, and evaluates 93 models from tiny open encoders to large commercial APIs. With bootstrap intervals, paired significance, and an Item Response Theory view of how sharply each task separates models, the suite cleanly orders most of the field into about a dozen tiers while leaving the top six statistically tied. An openly licensed, self-hostable model sits in that leading cluster, so frontier Portuguese quality does not require a paid API. A model’s rank on the global multilingual board predicts its Portuguese rank only moderately (Spearman ρ = 0.75), and one model that ranks third there falls near the bottom here, so a native benchmark measures something the multilingual boards do not.

Core claim

On 22 native Brazilian-Portuguese embedding tasks, multilingual leaderboard rank predicts Portuguese rank only moderately (Spearman ρ = 0.75 over 55 shared models; one model is 3rd there and 49th here), while an open self-hostable model reaches the statistically unresolved top tier of six, so strong Portuguese embedding quality does not require a commercial API and native evaluation is not redundant with global boards.

What carries the argument

MTEB-BR’s 22-task mean over native-only Portuguese tasks, backed by a statistical layer of per-task bootstrap confidence intervals, paired-bootstrap significance (tiers and probability-of-best), Item Response Theory task- and instance-level discrimination, and Borda robustness, which together turn the leaderboard into an uncertainty-aware measurement rather than a bare ranking.

Load-bearing premise

The filter that keeps only native Portuguese sources and then drops low-discrimination candidates yields a 22-task mean that truly reflects Portuguese embedding quality rather than legal-register skew, short-document bias under a 512-token cap, or memorized web text.

What would settle it

If, on a larger set of native Brazilian-Portuguese retrieval and dialogue tasks with long documents and no legal over-weighting, the same 55 models’ multilingual ranks correlated near 0.95 with the new mean and no open model reached the commercial frontier, the claim that native evaluation measures something distinct and that open models already match would fail.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Practitioners can pick Portuguese embedding models from native evidence and treat multilingual rank as a first filter only, not a proxy.
  • Within the top six, choice should turn on cost, license, latency, and context length rather than score, since the benchmark does not resolve their order.
  • Open Apache-2.0 models at the frontier remove per-token cost for self-hosting without sacrificing the quality tier measured here.
  • Retrieval tasks carry the sharpest separation signal; budgets for harder Portuguese evaluation are better spent there than on weak clustering or low-signal classification probes.
  • The same statistical layer (bootstrap tiers, paired significance, IRT discrimination) can be reused for other language-specific embedding suites.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If other languages show similar moderate cross-leaderboard correlations, native MTEB-style boards may be necessary wherever deployment text diverges from English-mined and translated retrieval.
  • The training-objective account of the Portuguese-encoder retrieval gap suggests fine-tuning strong multilingual or decoder backbones on native Portuguese retrieval data is a more direct path than more language-specific pretraining alone.
  • Instance-level IRT recovery of rankings from the most informative examples points toward cheaper ongoing leaderboard maintenance as new models appear.
  • Silent closed-API drift and short-document caps mean published commercial ranks will need periodic re-runs against the same native suite before production lock-in.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces MTEB-BR, a native Brazilian-Portuguese text-embedding benchmark of 22 tasks across seven MTEB categories, constructed under a strict native-source filter that excludes machine-translated corpora (notably mMARCO-PT). It evaluates 93 models (73 open-weight, 20 closed APIs; 23M–27B parameters) and reports a statistical layer: per-task bootstrap CIs, paired-bootstrap significance and tiers, task- and instance-level IRT discrimination, Borda robustness, and a cross-leaderboard Spearman correlation. Headline results are that the suite separates roughly a dozen model tiers while leaving the top six unresolved; an Apache-2.0 open model (Qwen3-Embedding-8B) sits in that frontier tier; and multilingual MTEB rank predicts Portuguese rank only moderately (ρ=0.75 over 55 shared models, with large inversions such as Llama-Embed-Nemotron-8B 3rd→49th). Code, tasks, results, and a public leaderboard are released.

Significance. The work fills a clear gap: Portuguese embedding evaluation has relied on translated or thinly covered multilingual suites. The native construction, large and mixed open/closed model panel, and especially the transferable statistical layer (bootstrap tiers, IRT discrimination, Borda check) are genuine contributions that other language-specific MTEB extensions can reuse. The open-model parity result and the moderate cross-leaderboard correlation are practically actionable and are backed by a released 93×22 score matrix with explicit uncertainty. Full open release of code (Apache-2.0), results (CC-BY 4.0), and an interactive leaderboard further strengthens the contribution.

major comments (3)
  1. [§X Limitations; §IX; Appendix A] §X and Appendix A: the uniform 512-token cap is acknowledged, and a single-model check on BR-TaxQA-R shows a large nDCG@10 lift when the cap is raised. Because retrieval is the dominant separator (§VIII, Table IV) and a frontier endpoint (voyage-context-4) is a long-context model, the cap conditions both the retrieval ranking and the cost–Pareto of §IX. A short multi-model ablation (or at least reporting scores at 512 vs. native context for the long-context APIs and open models that support it) would make the frontier claim more robust rather than leaving the penalty as a single-model bound.
  2. [§III; §VIII] §III exclusions and §VIII: the discrimination screen that drops low-separating candidates is correctly flagged as biasing reported a_t upward, and the authors note Spearman 0.99 before/after the final cut. For a benchmark paper whose design claim rests on native + discriminating tasks, a brief appendix table of the dropped candidates (name, category, reason, and if available a_t or score variance) would let readers judge survivor bias quantitatively rather than only via the authors’ qualitative screen description.
  3. [§VII] §VII: the cross-leaderboard ρ=0.75 is a central claim and is well supported by the rank-inversion example and category-wise breakdowns. The live HuggingFace board is a drifting, differently composed reference; the paper already notes this. A one-sentence sensitivity check (e.g., correlation restricted to models with stable identifiers, or against the published MMTEB headline ranking already mentioned) would further insulate the “native boards measure something different” conclusion from snapshot dependence.
minor comments (6)
  1. [Table III; §VI] Table III caption and §VI: score CIs are correctly described as marginal, not paired; consider adding a one-line pointer in the main text to Appendix F so readers do not misread overlapping CIs as non-significance.
  2. [Figure 1; §III] Figure 1 and the diversity diagnostic: mean-centered cosine is the right choice; stating the three-model panel and mean pairwise Spearman ρ=0.82 in the figure caption (not only the body) would make the plot self-contained.
  3. [Table II] Table II is dense (two-panel metadata for 93 models). A machine-readable companion (already released) is fine; in the PDF, a short family-level summary table or sorting by family would improve scannability.
  4. [§VIII-A; Abstract; §XI] §VIII-A instance-level IRT: the efficiency and misfit results are carefully scoped as a case study on four dichotomous tasks. Ensure the abstract and conclusion do not imply suite-wide adaptive evaluation without that qualifier.
  5. [Title page; References] Minor typography: “S ˜ao Paulo”, “Avaliac ¸˜ao”, and similar accent/spacing artifacts appear in affiliations and references; a pass with a Portuguese-aware PDF toolchain would clean them.
  6. [§III; Appendix D] Appendix D (mMARCO train–test overlap) is valuable supporting evidence for the native-source filter; a single sentence in §III pointing to the cohort rank-shift figure would help readers who skip appendices.

Circularity Check

0 steps flagged

No significant circularity: empirical rankings and cross-leaderboard correlation rest on held-out native scores, not on self-defined or fitted-as-prediction quantities.

full rationale

MTEB-BR is an empirical benchmark paper. Its load-bearing claims—the 22-task mean ranking of 93 models, the unresolved top-six frontier tier under paired bootstrap, the open-weight model (Qwen3-Embedding-8B) inside that tier, and the moderate Spearman ρ=0.75 vs. the multilingual MTEB board over 55 shared models—are computed directly from a released score matrix on natively sourced tasks. They do not reduce to parameters fitted to the same quantities they are said to predict. The IRT layer is explicitly a descriptive least-squares 2-PL fit on the observed continuous score matrix (task-level a_t tracks raw cross-model variance at ρ=0.88; abilities recover the mean ranking at Kendall τ=0.91); the authors do not present a_t or θ_m as first-principles predictions. Task selection used discrimination as a post-hoc screen and the authors state the resulting a_t values are biased upward and are reported only to document the screen; they also show the ranking is essentially unchanged after the cut (Spearman 0.99). Shared-corpus reformulations (ASSIN, Quati, JurisTCU) are acknowledged and a drop-redundant check leaves Kendall τ=0.97 with the full mean. There is no self-citation uniqueness theorem, no ansatz smuggled from prior author work as a forced form, and no renaming of a known result as a derived law. Limitations (register skew, 512-token cap, web-text contamination, convenience panel) are scope caveats, not circular reductions. Score 0; steps empty.

Axiom & Free-Parameter Ledger

4 free parameters · 4 axioms · 1 invented entities

As an empirical benchmark paper the load-bearing commitments are design choices and standard statistical machinery rather than free physical constants or new ontological entities. The free parameters are the operational decisions that fix the ranking (task count, token cap, bootstrap size, IRT fitting procedure). Axioms are the inherited MTEB metrics and the 2-PL response model. No new particles, forces or latent constructs beyond the ordinary IRT ability θ are postulated.

free parameters (4)
  • uniform max_seq_length = 512 tokens
    Hard-coded encoding cap applied to every model and task; the authors show on BR-TaxQA-R that lifting it to 2048 raises nDCG@10 by ~25 % for bge-m3, so the cap materially conditions both retrieval ranks and the cost-Pareto frontier.
  • number of bootstrap resamples = 10 000 (tasks) / 200 (IRT panel)
    Chosen by the authors; intervals and p-values depend on this Monte-Carlo size.
  • 2-PL least-squares fit (5 000 gradient steps, θ mean-centered)
    Continuous-score adaptation of the classic dichotomous IRT model; recovered a_t and θ_m are therefore relative to this fitting procedure and the 93-model panel.
  • final 22-task suite after discrimination screen
    Many candidate Portuguese tasks were dropped for degeneracy, low a_t or licensing; the reported discriminations are those of the survivors and are acknowledged to be biased upward.
axioms (4)
  • domain assumption MTEB per-category primary metrics (accuracy, subset accuracy, Spearman, V-measure, nDCG@10, MAP) are the correct evaluation targets for embedding quality.
    Inherited unchanged from Muennighoff et al. (2023) so that scores remain comparable; never re-derived.
  • domain assumption A two-parameter logistic response function adequately describes continuous normalized task scores as a function of latent model ability.
    Standard IRT modeling choice applied at task and instance level (§VIII); used only descriptively.
  • ad hoc to paper Machine-translated retrieval corpora (mMARCO-PT, mkqa-PT) introduce artifacts that can mask true model differences and are therefore excluded by construction.
    Central design axiom of §III; supported by the train-on-test analysis in Appendix D but still a modeling decision.
  • domain assumption Unweighted mean of the 22 primary metrics is a fair headline ranking (category-balanced mean yields Kendall τ=0.91).
    Follows MTEB convention; robustness checks are supplied.
invented entities (1)
  • MTEB-BR task suite (22 native formulations) independent evidence
    purpose: Provide the first consolidated native Brazilian-Portuguese embedding benchmark free of machine translation.
    The suite itself is the paper’s primary constructed object; several tasks (WikiCatClus, MedPT reformulations, SciELO/StackOverflow/JurisTCU clustering, FaqBacen, FaQuAD-IR) are newly packaged by the authors from public sources.

pith-pipeline@v1.1.0-grok45 · 35221 in / 3228 out tokens · 32127 ms · 2026-07-11T16:54:30.249232+00:00 · methodology

0 comments
read the original abstract

Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEB-BR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, and reranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, and a cross-leaderboard correlation. Three findings stand out. The benchmark cleanly separates about a dozen tiers of models, though the top six are statistically too close to order. An openly licensed, self-hostable model reaches that leading tier, so strong Portuguese embedding quality does not require a commercial API. And a model's rank on the global multilingual leaderboard predicts its Portuguese rank only moderately (Spearman rho = 0.75 over 55 shared models; one model ranks 3rd there and 49th here), so a native benchmark measures something the multilingual boards do not. We release every task, our code, and a public leaderboard, so practitioners can choose Portuguese embedding models on native evidence.

Figures

Figures reproduced from arXiv: 2607.04581 by Tardelli Ronan Coelho Stekel.

Figure 1
Figure 1. Figure 1: Corpus content overlap. Mean-centered pairwise cosine similarity between per-task mean embeddings (100 documents per task), averaged over three architecturally diverse models (multilingual-E5-base, Qwen3- Embedding-0.6B, Serafim-100m). Within-category similarity exceeds across￾category (values span [−0.45, 1.00]); the lone near-duplicate (1.00) is Ass￾inRTE versus AssinSTS, which share a source corpus. mul… view at source ↗
Figure 2
Figure 2. Figure 2: HuggingFace MTEB multilingual versus MTEB-BR rank, expressed [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Resource versus quality on MTEB-BR; both panels share the quality axis (22-task mean, [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Rank-shift on translated mMARCO-PT by training-data cohort (71 [PITH_FULL_IMAGE:figures/full_fig_p013_4.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

    cs.CL 2026-07 conditional novelty 5.5

    Portuguese embedding quality is task-dependent; multilingual leaderboards misrank models, while Portuguese contrastive fine-tuning plus MRL improves STS and retrieval under compact dimensions.

  2. Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders

    cs.CL 2026-07 conditional novelty 5.0

    MTEB-PT, a 14-task Portuguese slice of MMTEB, shows embedding rankings for Portuguese are task-dependent, multilingual ranks transfer only partially, and Portuguese MRL fine-tuning helps most on STS.

Reference graph

Works this paper leans on

68 extracted references · 1 canonical work pages · cited by 1 Pith paper

  1. [1]

    MTEB: Massive text embedding benchmark,

    N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive text embedding benchmark,” inProceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. Dubrovnik, Croatia: Association for Computational Linguistics, May 2023, pp. 2014–2037. [Online]. Available: https: //aclanthology.org/2023.eacl-main.148

  2. [2]

    MMTEB: Massive multilingual text embedding benchmark,

    K. Enevoldsen, I. Chunget al., “MMTEB: Massive multilingual text embedding benchmark,” inProceedings of the 13th International Con- ference on Learning Representations, 2025, arXiv:2502.13595

  3. [3]

    mMARCO: A multilingual version of the MS MARCO passage ranking dataset,

    L. Bonifacio, V . Jeronymo, H. Q. Abonizio, I. Campiotti, M. Fadaee, R. Lotufo, and R. Nogueira, “mMARCO: A multilingual version of the MS MARCO passage ranking dataset,” 2021

  4. [4]

    VN-MTEB: Vietnamese massive text embedding benchmark,

    L. Pham, T. Luu, T. V o, M. Nguyen, and V . Hoang, “VN-MTEB: Vietnamese massive text embedding benchmark,” 2025

  5. [5]

    C- Pack: Packed resources for general Chinese embeddings,

    S. Xiao, Z. Liu, P. Zhang, N. Muennighoff, D. Lian, and J.-Y . Nie, “C- Pack: Packed resources for general Chinese embeddings,” inProceed- ings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. New York, NY , USA: Associ- ation for Computing Machinery, 2024, pp. 641–649, arXiv:2309.07597

  6. [6]

    The Scandinavian embedding benchmarks: Comprehensive assessment of multilingual and monolingual text embedding,

    K. Enevoldsen, M. Kardos, N. Muennighoff, and K. L. Nielbo, “The Scandinavian embedding benchmarks: Comprehensive assessment of multilingual and monolingual text embedding,” inAdvances in Neural Information Processing Systems, vol. 37. Curran Associates, Inc., 2024, pp. 40 336–40 358, datasets and Benchmarks Track

  7. [7]

    MTEB-French: Resources for French sentence embedding evaluation and analysis,

    M. Ciancone, I. Kerboua, M. Schaeffer, and W. Siblini, “MTEB-French: Resources for French sentence embedding evaluation and analysis,” 2024

  8. [8]

    PL-MTEB: Polish massive text embedding benchmark,

    R. Po ´swiata, S. Dadas, and M. Perełkiewicz, “PL-MTEB: Polish massive text embedding benchmark,” 2024, accepted to Findings of ACL 2026

  9. [9]

    The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design,

    A. Snegirev, M. Tikhonova, A. Maksimova, A. Fenogenova, and A. Abramov, “The Russian-focused embedders’ exploration: ruMTEB benchmark and Russian embedding model design,” inProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Albuque...

  10. [10]

    FaMTEB: Massive text embedding benchmark in Persian language,

    E. Zinvandiet al., “FaMTEB: Massive text embedding benchmark in Persian language,” 2025, eMNLP 2025 Findings

  11. [11]

    MTEB-NL and E5-NL: Embedding benchmark and models for Dutch,

    N. Banaret al., “MTEB-NL and E5-NL: Embedding benchmark and models for Dutch,” 2025

  12. [12]

    German text embedding clustering benchmark,

    S. Wehrli, B. Arnrich, and C. Irrgang, “German text embedding clustering benchmark,” inProceedings of the 19th Conference on Natural Language Processing (KONVENS 2023). Ingolstadt, Germany: Association for Computational Linguistics, Sep. 2023, pp. 187–201. [Online]. Available: https://aclanthology.org/2023.konvens-main.20/

  13. [13]

    Ruri: Japanese general text embeddings,

    H. Tsukagoshi and R. Sasano, “Ruri: Japanese general text embeddings,” 2024

  14. [14]

    SentEval: An evaluation toolkit for universal sentence representations,

    A. Conneau and D. Kiela, “SentEval: An evaluation toolkit for universal sentence representations,” inProceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018). Miyazaki, Japan: European Language Resources Association, 2018. [Online]. Available: https://aclanthology.org/L18-1269

  15. [15]

    BEIR: A heterogenous benchmark for zero-shot evaluation of infor- mation retrieval models,

    N. Thakur, N. Reimers, A. R ¨uckl´e, A. Srivastava, and I. Gurevych, “BEIR: A heterogenous benchmark for zero-shot evaluation of infor- mation retrieval models,” inProceedings of the NeurIPS 2021 Track on Datasets and Benchmarks, 2021, arXiv:2104.08663

  16. [16]

    MIRACL: A multilingual retrieval dataset covering 18 diverse languages,

    X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, and J. Lin, “MIRACL: A multilingual retrieval dataset covering 18 diverse languages,”Transactions of the Association for Computational Linguistics, vol. 11, pp. 1114–1131,

  17. [17]

    Available: https://aclanthology.org/2023.tacl-1.63/

    [Online]. Available: https://aclanthology.org/2023.tacl-1.63/

  18. [18]

    BEIR-PL: Zero shot information retrieval benchmark for the Polish language,

    K. Wojtasik, V . Shishkin, K. Wołowiec, A. Janz, and M. Piasecki, “BEIR-PL: Zero shot information retrieval benchmark for the Polish language,” 2024

  19. [19]

    Napolab: The natural portuguese language bench- mark,

    R. C. Rodrigues, “Napolab: The natural portuguese language bench- mark,” https://github.com/ruanchaves/napolab, 2023

  20. [20]

    PORTULAN ExtraGLUE datasets and models: Kick-starting a benchmark for the neural processing of Portuguese,

    T. F. Os ´orio, B. Leite, H. Lopes Cardoso, L. Gomes, J. Rodrigues, R. Santos, and A. Branco, “PORTULAN ExtraGLUE datasets and models: Kick-starting a benchmark for the neural processing of Portuguese,” inProceedings of the 17th Workshop on Building and Using Comparable Corpora (BUCC) @ LREC-COLING 2024, P. Zweigenbaum, R. Rapp, and S. Sharoff, Eds. Torin...

  21. [21]

    Beyond multilingual averages: MTEB-PT, a benchmark for Portuguese sentence encoders,

    L. H. T. Okamura, A. Alcoforado, and A. H. Reali Costa, “Beyond multilingual averages: MTEB-PT, a benchmark for Portuguese sentence encoders,” 2026, accepted at BRACIS 2026

  22. [22]

    HateBR: A large expert annotated corpus of Brazilian Instagram comments for offensive language and hate speech detection,

    F. Vargas, I. Carvalho, F. Rodrigues de G ´oes, T. Pardo, and F. Benevenuto, “HateBR: A large expert annotated corpus of Brazilian Instagram comments for offensive language and hate speech detection,” inProceedings of the Thirteenth Language Resources and Evaluation Conference. Marseille, France: European Language Resources Association, Jun. 2022, pp. 717...

  23. [23]

    FACTCK.BR: a new dataset to study fake news,

    J. Moreno and G. Bressan, “FACTCK.BR: a new dataset to study fake news,” inProceedings of the 25th Brazillian Symposium on Multimedia and the Web (WebMedia ’19). New York, NY , USA: Association for Computing Machinery, 2019, pp. 525–527

  24. [24]

    ToxSyn-PT: A synthetic fine-grained dataset of minority- targeted toxic language in Portuguese,

    I. A. Brito, J. S. Dollis, F. B. Farber, D. Fernandes, and A. R. Galv˜ao Filho, “ToxSyn-PT: A synthetic fine-grained dataset of minority- targeted toxic language in Portuguese,” 2026

  25. [25]

    Vis˜ao Geral da Avaliac ¸˜ao de Similaridade Sem ˆantica e Infer ˆencia Textual,

    E. R. Fonseca, L. B. d. Santos, M. Criscuolo, and S. M. Alu ´ısio, “Vis˜ao Geral da Avaliac ¸˜ao de Similaridade Sem ˆantica e Infer ˆencia Textual,” Linguam´atica, vol. 8, no. 2, pp. 3–13, 2016. [Online]. Available: https: //www.linguamatica.com/index.php/linguamatica/article/view/v8n2-1

  26. [26]

    InferBR: A natural language inference dataset in Portuguese,

    L. Bencke, F. V . Pereira, M. K. Santos, and V . Moreira, “InferBR: A natural language inference dataset in Portuguese,” inProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). Torino, Italia: ELRA and ICCL, May 2024, pp. 9050–9060. [Online]. Available: https://aclantholo...

  27. [27]

    The ASSIN 2 shared task: A quick overview,

    L. Real, E. Fonseca, and H. Gonc ¸alo Oliveira, “The ASSIN 2 shared task: A quick overview,” inComputational Processing of the Portuguese Language (PROPOR 2020), ser. Lecture Notes in Computer Science, vol. 12037. Cham: Springer, 2020, pp. 406–412

  28. [28]

    MedPT: A massive medical question answering dataset for Brazilian-Portuguese speakers,

    F. B. F ¨arber, I. A. Brito, J. S. Dollis, P. S. F. B. Ribeiro, R. T. Sousa, and A. R. Galv ˜ao Filho, “MedPT: A massive medical question answering dataset for Brazilian-Portuguese speakers,” 2025, accepted at LREC 2026

  29. [29]

    Brazilian Portuguese Wikipedia categories clustering dataset,

    Wikimedia Foundation, “Brazilian Portuguese Wikipedia categories clustering dataset,” HuggingFace dataset mteb-br/wikipedia-categories, 2026, constructed from Brazilian- Portuguese Wikipedia (CC-BY-SA-3.0). [Online]. Available: https://huggingface.co/datasets/mteb-br/wikipedia-categories

  30. [30]

    JurisTCU: A Brazilian Portuguese information retrieval dataset with query relevance judgments,

    L. C. Fernandes, L. d. S. Ribeiro, M. V . B. de Castro, L. A. d. S. Pacheco, and E. F. d. O. Sandes, “JurisTCU: A Brazilian Portuguese information retrieval dataset with query relevance judgments,”Language Resources and Evaluation, 2026, also available as arXiv:2503.08379

  31. [31]

    SciELO Brazil scientific abstracts clustering dataset,

    SciELO Brazil, “SciELO Brazil scientific abstracts clustering dataset,” HuggingFace datasetmteb-br/scielo-clustering, 2026, constructed from the SciELO Brazil open-access library. [Online]. Available: https://huggingface.co/datasets/mteb-br/scielo-clustering

  32. [32]

    Portuguese Stack Overflow clustering dataset,

    Stack Exchange, “Portuguese Stack Overflow clustering dataset,” HuggingFace datasetmteb-br/stackoverflow-clustering, 2026, constructed from pt.stackoverflow.com question titles (CC-BY- SA). [Online]. Available: https://huggingface.co/datasets/mteb-br/ stackoverflow-clustering

  33. [33]

    FaQuAD: Reading comprehension dataset in the domain of Brazilian higher education,

    H. F. Sayama, A. V . Araujo, and E. R. Fernandes, “FaQuAD: Reading comprehension dataset in the domain of Brazilian higher education,” pp. 443–448, 2019. [Online]. Available: https://doi.org/10.1109/BRACIS. 2019.00084

  34. [34]

    Quati: A Brazilian Portuguese information retrieval dataset from native speakers,

    E. d. Oliveira, M. Bueno, R. Nogueira, R. Lotufo, and J. Pereira, “Quati: A Brazilian Portuguese information retrieval dataset from native speakers,” inProceedings of the 15th Brazilian Symposium in Information and Human Language Technology, D. B. Claro and A. Pagano, Eds. Bel ´em do Par ´a, Brazil: Association for Computational Linguistics, 2024, pp. 185...

  35. [35]

    FaqBacen: Banco central do brasil public FAQ retrieval dataset,

    Banco Central do Brasil, “FaqBacen: Banco central do brasil public FAQ retrieval dataset,” HuggingFace datasetmteb-br/faq-bacen, 2026, reformulated from the Banco Central do Brasil public FAQ. [Online]. Available: https://huggingface.co/datasets/mteb-br/faq-bacen 16

  36. [36]

    BR-TaxQA-R: A dataset for question answering with references for Brazilian personal income tax law, including case law,

    J. Domingos J ´unior, A. Faria, E. Seiti de Oliveira, E. de Brito, M. Teoto- nio, A. Assumpc ¸˜ao, D. Carmo, R. Lotufo, and J. Pereira, “BR-TaxQA-R: A dataset for question answering with references for Brazilian personal income tax law, including case law,” 2025

  37. [37]

    RoBERTaLexPT: A legal RoBERTa model pretrained with deduplication for Portuguese,

    E. A. S. Garcia, N. F. F. Silva, F. Siqueira, J. R. Gomes, H. O. Albuquerque, E. Souza, E. Lima, and A. de Carvalho, “RoBERTaLexPT: A legal RoBERTa model pretrained with deduplication for Portuguese,” inProceedings of the 16th International Conference on Computational Processing of Portuguese (PROPOR 2024). Santiago de Compostela, Galicia/Spain: Associati...

  38. [38]

    BRIGHTER: BRIdging the gap in human-annotated textual emotion recognition datasets for 28 languages,

    S. H. Muhammad, N. Ousidhoum, I. Abdulmumin, J. P. Wahle, T. Ruas et al., “BRIGHTER: BRIdging the gap in human-annotated textual emotion recognition datasets for 28 languages,” inProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computa- tional Linguistics, Jul...

  39. [39]

    Qwen3 embedding: Advancing text embedding and reranking through foundation models,

    Y . Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou, “Qwen3 embedding: Advancing text embedding and reranking through foundation models,” 2025

  40. [40]

    Octen-embedding model family,

    Octen, “Octen-embedding model family,” Hugging Face model card, 2025. [Online]. Available: https://huggingface.co/Octen/ Octen-Embedding-8B

  41. [41]

    F2LLM-v2: Inclusive, performant, and efficient embeddings for a multilingual world,

    Z. Zhang, Z. Liao, H. Yu, P. Di, and R. Wang, “F2LLM-v2: Inclusive, performant, and efficient embeddings for a multilingual world,” 2026

  42. [42]

    Harrier-OSS-v1 embedding family,

    Microsoft, “Harrier-OSS-v1 embedding family,” Hugging Face model card, 2026. [Online]. Available: https://huggingface.co/microsoft/ harrier-oss-v1-27b

  43. [43]

    SFR-embedding-mistral: Enhance text retrieval with transfer learning,

    R. Meng, Y . Liu, S. R. Joty, C. Xiong, Y . Zhou, and S. Yavuz, “SFR-embedding-mistral: Enhance text retrieval with transfer learning,” Salesforce AI Research Blog, 2024. [Online]. Available: https: //www.salesforce.com/blog/sfr-embedding/

  44. [44]

    Towards general text embeddings with multi-stage contrastive learning,

    Z. Li, X. Zhang, Y . Zhang, D. Long, P. Xie, and M. Zhang, “Towards general text embeddings with multi-stage contrastive learning,”arXiv preprint arXiv:2308.03281, 2023

  45. [45]

    Linq-Embed-Mistral technical report,

    C. Choi, J. Kim, S. Lee, J. Kwon, S. Gu, Y . Kim, M. Cho, and J.-y. Sohn, “Linq-Embed-Mistral technical report,” 2024

  46. [46]

    KaLM-Embedding-V2: Superior training techniques and data inspire a versatile embedding model,

    X. Zhao, X. Hu, Z. Shan, S. Huang, Y . Zhou, X. Zhang, Z. Sun, Z. Liu, D. Li, X. Wei, Y . Pan, Y . Xiang, M. Zhang, H. Wang, J. Yu, B. Hu, and M. Zhang, “KaLM-Embedding-V2: Superior training techniques and data inspire a versatile embedding model,” 2025

  47. [47]

    Llama-Embed-Nemotron-8B: A universal text embedding model for multilingual and cross-lingual tasks,

    Y . Babakhin, R. Osmulski, R. Ak, G. Moreira, M. Xu, B. Schifferer, B. Liu, and E. Oldridge, “Llama-Embed-Nemotron-8B: A universal text embedding model for multilingual and cross-lingual tasks,” 2025

  48. [48]

    BidirLM: From text to omnimodal bidirectional encoders by adapting and composing causal LLMs,

    N. Boizard, T. Deschamps-Berger, H. Gisserot-Boukhlef, C. Hudelot, and P. Colombo, “BidirLM: From text to omnimodal bidirectional encoders by adapting and composing causal LLMs,” 2026

  49. [49]

    jua: Domain-adaptive dense retrieval embeddings for Brazilian legal search,

    J. Pereira, R. Lotufo, and L. Bonifacio, “jua: Domain-adaptive dense retrieval embeddings for Brazilian legal search,” HuggingFace modelufca-llms/jua-4B-mixed, 2026. [Online]. Available: https: //huggingface.co/ufca-llms/jua-4B-mixed

  50. [50]

    Multilingual E5 text embeddings: A technical report,

    L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei, “Multilingual E5 text embeddings: A technical report,” 2024

  51. [51]

    Sentence-BERT: Sentence embeddings using Siamese BERT-networks,

    N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” inProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China: Association for Computational Linguistics, Nov. 2019, pp. 3982–399...

  52. [52]

    Granite embedding multilingual R2 models,

    P. Awasthy, A. Trivedi, Y . Yang, K. Barker, Y . Li, B. Iyer, M. Franz, J. Bross, M. Doshi, P. Vignesh, V . Kumar, T. Ward, A. Daniels, M. Lee, L. Lastras, J. Sen, and R. Florian, “Granite embedding multilingual R2 models,” 2026

  53. [53]

    jina-embeddings-v5-text: Task-targeted embed- ding distillation,

    M. K. Akram, S. Sturua, N. Havriushenko, Q. Herreros, M. G ¨unther, M. Werk, and H. Xiao, “jina-embeddings-v5-text: Task-targeted embed- ding distillation,” 2026

  54. [54]

    M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation,

    J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu, “M3-embedding: Multi-linguality, multi-functionality, multi-granularity text embeddings through self-knowledge distillation,” inFindings of the Association for Computational Linguistics: ACL 2024. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 2318–2335. [Online]. Avail...

  55. [55]

    Arctic-embed 2.0: Multi- lingual retrieval without compromise,

    P. Yu, L. Merrick, G. Nuti, and D. Campos, “Arctic-embed 2.0: Multi- lingual retrieval without compromise,” 2024

  56. [56]

    Open sentence embeddings for Portuguese with the Serafim PT* encoders family,

    L. Gomes, A. Branco, J. Silva, J. Rodrigues, and R. Santos, “Open sentence embeddings for Portuguese with the Serafim PT* encoders family,” inProgress in Artificial Intelligence – 23rd EPIA Conference on Artificial Intelligence, EPIA 2024, M. F. Santos, J. Machado, P. Novais, P. Cortez, and P. M. Moreira, Eds. Cham: Springer, 2025, pp. 267–279, arXiv:2407.19527

  57. [57]

    Fostering the ecosystem of open neural encoders for Portuguese with Albertina PT* family,

    R. Santos, J. Rodrigues, L. Gomes, J. Silva, A. Branco, H. Lopes Car- doso, T. Freitas Os ´orio, and B. Leite, “Fostering the ecosystem of open neural encoders for Portuguese with Albertina PT* family,” 2024

  58. [58]

    BERTimbau: Pretrained BERT models for Brazilian Portuguese,

    F. Souza, R. Nogueira, and R. Lotufo, “BERTimbau: Pretrained BERT models for Brazilian Portuguese,” inProceedings of the 9th Brazilian Conference on Intelligent Systems (BRACIS 2020). Springer, 2020, pp. 403–417

  59. [59]

    A semantic search system for the Supremo Tribunal de Justic ¸a,

    R. Melo, P. A. Santos, and J. Dias, “A semantic search system for the Supremo Tribunal de Justic ¸a,” inProgress in Artificial Intelligence – 22nd EPIA Conference on Artificial Intelligence (EPIA 2023), Pro- ceedings, Part II, ser. Lecture Notes in Computer Science, vol. 14116. Springer Nature Switzerland, 2023, pp. 142–154

  60. [60]

    Language- agnostic BERT sentence embedding,

    F. Feng, Y . Yang, D. Cer, N. Arivazhagan, and W. Wang, “Language- agnostic BERT sentence embedding,” inProceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 2022, pp. 878–891. [Online]. Available: https://aclanthology.org/2022.acl-long.62

  61. [61]

    EmbeddingGemma: Powerful and lightweight text representations,

    H. Schechter Vera, S. Dua, B. Zhang, D. Salz, R. Mullins, S. Raghu- ram Panyam, S. Smoot, I. Naim, J. Zou, F. Chen, D. Ceret al., “EmbeddingGemma: Powerful and lightweight text representations,” 2025

  62. [62]

    PIXIE-rune-v1.0,

    Telepix, “PIXIE-rune-v1.0,” Hugging Face model card, 2025. [Online]. Available: https://huggingface.co/telepix/PIXIE-Rune-v1.0

  63. [63]

    PwC-embedding expr,

    SamilPwC AXNode-GenAI Lab, “PwC-embedding expr,” Hugging Face model card, 2025. [Online]. Available: https://huggingface.co/ SamilPwC-AXNode-GenAI/PwC-Embedding expr

  64. [64]

    What are the best Systems? New Perspectives on NLP Benchmarking,

    P. Colombo, N. Noiry, E. Irurozki, and S. Cl ´emenc ¸on, “What are the best Systems? New Perspectives on NLP Benchmarking,” inAdvances in Neural Information Processing Systems, vol. 35. Curran Associates, Inc., 2022, pp. 26 915–26 932, arXiv:2202.03799

  65. [65]

    R. K. Hambleton, H. Swaminathan, and H. J. Rogers,Fundamentals of Item Response Theory. Newbury Park, CA: Sage Publications, 1991

  66. [66]

    Building an evaluation scale using item response theory,

    J. P. Lalor, H. Wu, and H. Yu, “Building an evaluation scale using item response theory,” inProceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Austin, Texas: Association for Computational Linguistics, 2016, pp. 648–657. [Online]. Available: https://aclanthology.org/D16-1062/

  67. [67]

    tinyBenchmarks: evaluating LLMs with fewer examples,

    F. Maia Polo, L. Weber, L. Choshen, Y . Sun, G. Xu, and M. Yurochkin, “tinyBenchmarks: evaluating LLMs with fewer examples,” inProceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 235. PMLR, 2024, pp. 34 303–34 326. [Online]. Available: https://proceedings.mlr.press/v235/maia-polo24a.html

  68. [68]

    Evaluation examples are not equally informative: How should that change NLP leaderboards?

    P. Rodriguez, J. Barrow, A. Hoyle, J. P. Lalor, R. Jia, and J. Boyd-Graber, “Evaluation examples are not equally informative: How should that change NLP leaderboards?” inProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)....