Pith. sign in

REVIEW 3 major objections 4 minor 85 references

Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Prompt-based text embeddings stay accurate after cutting to 25% of their dimensions, and often far less for classification and clustering.

desk verdict Solid truncation-robustness result for classification/clustering embeddings, but the abstract overstates clustering and the explanatory ID/IsoScore analysis is built on a single Wikipedia sample. read the letter →

arxiv 2506.01435 v1 pith:CLZN363S submitted 2025-06-02 cs.CL

classification cs.CL
keywords textembeddingsprompt-baseddimensionalityreductionintrinsicisotropyembeddingredundancyMTEBsemantictextualsimilarity
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that the thousands of dimensions produced by prompt-based text embedding models are largely redundant: simply keeping the first 25% of the coordinates of each embedding leaves performance almost unchanged across classification, clustering, retrieval, and semantic similarity tasks. For classification and clustering, the paper reports that even reducing embeddings to less than 0.5% of their original dimensionality causes only small performance losses, and that a truncated large model can beat the full embeddings of a smaller model. The redundancy is task-dependent, and the paper explains the difference quantitatively: embeddings generated with classification or clustering prompts have lower intrinsic dimension and are less isotropic, while embeddings for retrieval and STS have higher intrinsic dimension and are more isotropic. If true, this means expensive prompt-based embeddings can be aggressively compressed for storage and computation without retraining, at least for some tasks.

What carries the argument

The argument is carried by three tools. The first is post-hoc truncation: embeddings are reduced by simply taking their first $d\in\mathbb{Z}_{>0}$ coordinates, without normalization, training, or learned projections; the paper checks that random coordinate selection and PCA give the same broad trends. The second is TwoNN, an estimator that infers intrinsic dimension from the ratios of distances to each point's two nearest neighbors, assumed to follow a Pareto distribution on a low-dimensional manifold. The third is IsoScore, computed from the normalized covariance matrix of a sample of embeddings, which measures how uniformly the embedding space is used, with values near 1 meaning isotropic and near 0 meaning anisotropic. Together these measures let the paper connect task-prompt choice to redundancy: classification and clustering prompts yield low-ID, anisotropic, highly redundant representations, while retrieval and STS prompts yield higher-ID, more isotropic, less redundant ones.

What would settle it

Recompute TwoNN and IsoScore on multiple independent samples from different domains and languages, or on long documents versus short sentences, and check whether classification and clustering prompts consistently show lower intrinsic dimensionality and lower isotropy than retrieval and STS prompts; if the ordering reverses or the gap vanishes, the geometric explanation for truncation robustness fails. A second check: if retrieval task performance stops declining when a different coordinate subset is kept, then the first-$d$ truncation result is an artifact of coordinate ordering rather than redundancy.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that the vector spaces produced by prompt-based models such as gte-Qwen2, E5-mistral, SFR-2, and mE5-large-inst contain far less information than their nominal dimensionality suggests, and that the amount of redundancy is controlled by the prompt. Using a simple truncation that keeps the first $d$ coordinates, average classification performance across five MTEB datasets stays nearly flat down to about 8 dimensions, and a 2-dimensional gte-Qwen2 embedding (76.34) outperforms the full 1024-dimensional E5-large embedding (75.69). Clustering degrades noticeably but remains strong down to about 128 dimensions for instruction-based models, whereas retrieval and STS deteriorate quickly as dimensions shrink. Measuring redundancy with TwoNN intrinsic dimension and IsoScore on 10,000 sampled Wikipedia paragraphs, the paper finds that classification and clustering prompts produce lower intrinsic dimensionality and lower isotropy, while retrieval and STS prompts produce higher values on both measures, matching the observed robustness ordering.

Load-bearing premise

The explanatory claim rests on the assumption that TwoNN intrinsic dimension and IsoScore, computed once on 10,000 sampled English Wikipedia paragraphs, are stable and valid enough to compare prompt types; the paper gives no repeated samples or error bars for these estimates.

Editorial extensions

If this is right

  • Storing and indexing prompt-based embeddings can be made much cheaper: keeping only the first quarter of coordinates preserves near-full performance across all four task families tested.
  • For classification and clustering, instruction-based models keep most of their performance at 8 to 128 dimensions (0.2 percent to 4 percent of the original), allowing large speedups in downstream linear classifiers and clustering pipelines.
  • A truncated embedding from a large prompt-based model can beat the full-dimensional embedding of a smaller dedicated model, so dimensionality rather than model size can be the deciding resource in some settings.
  • Task prompts are not just instructions; they implicitly set the geometric properties of the embedding space, with retrieval and STS requiring the fuller, more isotropic space.
  • Retrieval and STS systems should not aggressively truncate these embeddings, since performance falls quickly below roughly 25 percent of the original dimensions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the prompt controls redundancy, then prompt engineering could be used deliberately: rewriting a classification prompt to lower intrinsic dimensionality further may yield even more compressible embeddings without retraining.
  • The coordinate-order dependence of truncation suggests the first coordinates carry the task-relevant signal; inspecting which input features activate those coordinates could reveal what the model treats as class-defining.
  • The reported intrinsic dimension and IsoScore values come from one 10,000-text English Wikipedia sample; generalizing to other domains, text lengths, and languages is a direct testable extension that the authors flag as future work.
  • The finding that naive truncation competes with PCA and nonlinear methods implies that linear projection methods may be unnecessary for these models, while a learned per-task rotation might do even better.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper studies whether prompt-based text embedding models produce embeddings that can be truncated without retraining. Using MTEB classification, clustering, retrieval, and STS tasks across eight models, the authors report that simply keeping the first d dimensions preserves classification performance surprisingly well and that instruction-based models tolerate drastic truncation, while retrieval and STS degrade much more steeply. To explain this, they estimate intrinsic dimensionality with TwoNN and isotropy with IsoScore on a single 10,000-text English Wikipedia sample, finding that classification and clustering prompts yield lower intrinsic dimensionality and lower isotropy than retrieval and STS prompts. The paper includes additional experiments with random dimension selection, PCA, UMAP, and Isomap as controls and discusses related work on embedding dimensionality and isotropy.

Significance. If the truncation findings hold, the paper has clear practical value: for classification and some clustering workloads, large instruction-tuned embedding models could be deployed with substantially reduced storage and distance-computation costs without fine-tuning or specialized training objectives. The study is also valuable for documenting that prompt type systematically changes geometric properties of embeddings, and it connects this to downstream truncation robustness in a falsifiable way. Strengths of the manuscript include direct evaluation across multiple public models and task families, careful controls against coordinate-alignment artifacts via random dimension selection and PCA comparisons, and the use of standard public benchmarks and estimators. The main weakness is that the explanatory component rests on unreplicated measurements on a single corpus, and the abstract overgeneralizes the clustering result.

major comments (3)
  1. [Abstract; Section 2.4] The Abstract states that for classification and clustering, reduction to less than 0.5% of the original dimensionality yields very small degradation, but Section 2.4 does not support this for clustering. The text explicitly says clustering degradation is 'relatively noticeable' and reports that E5-large loses about 13 points when reduced to 128 dimensions (12.5% of its original 1024 dimensions). Even for gte-Qwen2, the robustness example is 128 dimensions, which is 3.6% of the original dimensionality, not 0.5%. The sub-0.5% claim is defensible for some instruction-based classification models (gte-Qwen2 and SFR-2 at 8 dimensions), but the abstract and conclusion should scope the claim to those models and tasks, or report the full set of results for clustering at the 0.5% cutoff.
  2. [Section 3.1; Section 3.2; Limitations] The explanatory claim that low intrinsic dimensionality and low isotropy cause truncation robustness is supported only by estimates computed on a single random sample of 10,000 English Wikipedia paragraphs (Section 3.1). Tables 1 and 2 report one ID and one IsoScore value per model/prompt condition, with no error bars, repeated samples, or hold-out validation, and the Limitations section explicitly concedes that text length, domain, and language variation were not analyzed. This is load-bearing because the MTEB task inputs used in Section 2 (Amazon reviews, Reddit titles, MS MARCO passages, STS pairs) differ substantially from Wikipedia paragraphs in genre, length, and label structure. To make the proposed mechanism convincing, the authors should either measure ID and IsoScore on the actual task corpora and show that the Table 1 ordering is preserved, or report variability across multiple independent samples and subsample sizes. Without this, the observed correlation could be an artifact of corpus choice, and the robustness could alternatively be explained by the logistic-regression classifier or clustering algorithm adapting to low-dimensional inputs rather than by redundancy inherent to the original embeddings.
  3. [Section 3.1; Tables 1 and 2] The reported ID and IsoScore values in Tables 1 and 2 are given without any uncertainty quantification. TwoNN estimates are known to depend on sample size and local density, and the paper itself notes that IsoScore can be unstable on small datasets and cites IsoScore* as a stabilized variant (footnote 7); the decision not to use IsoScore* is understandable because no training is involved, but a finite-sample bias check is still needed. Since the paper's key ordering claim is based on roughly tenfold differences in IsoScore and differences of about 10-40 in intrinsic dimension across prompt types, the authors should provide confidence intervals, bootstrap replicates, or sensitivity analyses over sample sizes to show that these differences are not within estimation noise.
minor comments (4)
  1. [Figure 6 caption] The caption of Figure 6 contains Japanese text ('分類タスク クラスタリング 検索文章 検索クエリ STS dim'); this should be translated to English for consistency with the rest of the manuscript.
  2. [Section 2.2] In the description of nomic, the phrase 'clustering: for clustering tasks' is missing a closing quotation mark after 'clustering:'; also, the sentence beginning 'clustering: for clustering tasks' should be a complete sentence or a properly formatted list item.
  3. [Section 4.1] The sentence 'Our study is the first to qualitatively examine what embeddings prompt-based text embedding models produce' appears to mean 'quantitatively examine' and also makes a strong novelty claim; I suggest rewording and softening the claim or providing explicit citation support for it.
  4. [Appendix C; reproducibility] The paper does not include a reproducibility statement or link to code; providing the evaluation scripts, the exact random seed and dimension-selection procedure for the Random method, and the preprocessing steps for the Wikipedia sample would make the experiments easier to verify.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: truncation robustness is directly evaluated, and the ID/isotropy analysis is an independent empirical correlate, not a derived prediction.

full rationale

The paper's central claim that prompt-based embeddings are highly redundant is supported by direct evaluations of downstream task performance after dimensionality reduction. Classification, clustering, retrieval, and STS scores are measured on MTEB datasets with no parameter fitted from the target result and then renamed as a prediction; the robustness curves are empirical outputs, not constructed from the redundancy measures. The explanatory analysis in Section 3 uses two external estimators, TwoNN for intrinsic dimensionality and IsoScore for isotropy, computed on an independently sampled set of 10,000 English Wikipedia paragraphs. The connection between low ID/low isotropy and truncation robustness is presented as a correlation across task types, not as a mathematical derivation in which the robustness result is presupposed by the estimators. The only self-citation, Tsukagoshi et al. (2021), is invoked for the standard use of Spearman's rank correlation in STS evaluation and is not load-bearing for any central claim. The Limitations section explicitly concedes that underlying causal factors are not clarified and that variation across text length, domain, and language was not analyzed; this is a scope and validity concern, not evidence of circularity. No equation or definition in the paper reduces a predicted quantity to an input quantity by construction, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is empirical; its central claim rests on the validity of the chosen estimators (TwoNN, IsoScore), the representativeness of the MTEB subsets and the single Wikipedia sample, and the public behavior of the evaluated model checkpoints. There are no fitted free parameters and no invented entities.

assumptions (4)
  • domain assumption TwoNN intrinsic dimension estimates are valid and comparable across prompt types for 10,000 sampled texts.
    Section 3.1 relies on TwoNN's manifold assumption, that nearest-neighbor distance ratios follow a Pareto distribution, to compare IDs across prompts; no error bars or validation on synthetic data are provided.
  • domain assumption IsoScore provides stable isotropy measurements for these embedding sets.
    Section 3.1 applies original IsoScore, citing reliability when sample size exceeds dimensionality; stability is not demonstrated for the anisotropic classification and clustering embeddings with scores near 0.005.
  • domain assumption Downsampled MTEB retrieval sets and the selected classification and clustering datasets are representative of those task families.
    Section 2.1 uses MTEB's official downsampled retrieval sets with 250 hard negatives and a maximum of 1,000 examples, plus a subset of classification and clustering datasets; generalization to other benchmarks is assumed.
  • domain assumption The 10,000 English Wikipedia paragraphs sampled for ID and IsoScore reflect the text distribution relevant to the downstream tasks.
    Section 3.1 samples a single Wikipedia dump; the authors themselves note in Limitations that domain and language variation are not analyzed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings." pith.science (2026). https://pith.science/paper/CLZN363S

@misc{pith2026250601435,
  author       = {Pith},
  title        = {Pith review of: Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CLZN363S}},
  note         = {Machine review of arXiv:2506.01435}
}
read the original abstract

Prompt-based text embedding models, which generate task-specific embeddings upon receiving tailored prompts, have recently demonstrated remarkable performance. However, their resulting embeddings often have thousands of dimensions, leading to high storage costs and increased computational costs of embedding-based operations. In this paper, we investigate how post-hoc dimensionality reduction applied to the embeddings affects the performance of various tasks that leverage these embeddings, specifically classification, clustering, retrieval, and semantic textual similarity (STS) tasks. Our experiments show that even a naive dimensionality reduction, which keeps only the first 25% of the dimensions of the embeddings, results in a very slight performance degradation, indicating that these embeddings are highly redundant. Notably, for classification and clustering, even when embeddings are reduced to less than 0.5% of the original dimensionality the performance degradation is very small. To quantitatively analyze this redundancy, we perform an analysis based on the intrinsic dimensionality and isotropy of the embeddings. Our analysis reveals that embeddings for classification and clustering, which are considered to have very high dimensional redundancy, exhibit lower intrinsic dimensionality and less isotropy compared with those for retrieval and STS.

Figures

Figures reproduced from arXiv: 2506.01435 by the authors.

Figure 6
Figure 6. ID under dimensionality reduction. The dashed line represents the actual dimensions. ing tasks, while yielding embeddings with lower redundancy for retrieval and STS tasks. ID and Isotropy with Dimensionality Reduc￾tion As in Section 2, we performed dimensional￾ity reduction on the embeddings and evaluated the changes in intrinsic dimension and isotropy. We measured the ID and IsoScore at each dimension us￾ing embed… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

85 extracted references · 20 canonical work pages

  1. [1]

    Williams

    Hervé Abdi and Lynne J. Williams. 2010. https://doi.org/10.1002/wics.101 Principal component analysis . WIREs Computational Statistics, 2(4):433--459

  2. [2]

    Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, I \ n igo Lopez-Gazpio, Montse Maritxalar, Rada Mihalcea, German Rigau, Larraitz Uria, and Janyce Wiebe. 2015. https://doi.org/10.18653/v1/S15-2045 SemEval-2015 Task 2: Semantic Textual Similarity, English, Spanish and Pilot on Interpretability . In Procee...

  3. [3]

    Eneko Agirre, Carmen Banea, Claire Cardie, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Weiwei Guo, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2014. https://doi.org/10.3115/v1/S14-2010 SemEval-2014 Task 10: Multilingual Semantic Textual Similarity . In Proceedings of the 8th International Workshop on Semantic Evaluation ( S em E val) , pages 81--91

  4. [4]

    Eneko Agirre, Carmen Banea, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, Rada Mihalcea, German Rigau, and Janyce Wiebe. 2016. https://doi.org/10.18653/v1/S16-1081 SemEval-2016 Task 1: Semantic Textual Similarity, Monolingual and Cross-Lingual Evaluation . In Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval), pages 497--511

  5. [5]

    Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012. https://www.aclweb.org/anthology/S12-1051 SemEval-2012 Task 6: A Pilot on Semantic Textual Similarity . In * SEM 2012: The First Joint Conference on Lexical and Computational Semantics -- Semantic Evaluation ( S em E val) , pages 385--393

  6. [6]

    Eneko Agirre, Daniel Cer, Mona Diab, Aitor Gonzalez-Agirre, and Weiwei Guo. 2013. https://www.aclweb.org/anthology/S13-1004 *SEM 2013 shared task: Semantic Textual Similarity . In Second Joint Conference on Lexical and Computational Semantics (* SEM ) , pages 32--43

  7. [7]

    Mira Ait-Saada and Mohamed Nadif. 2023. https://doi.org/10.18653/v1/2023.acl-short.103 Is Anisotropy Truly Harmful? A Case Study on Text Clustering . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), pages 1194--1203

  8. [8]

    Macke, and Davide Zoccolan

    Alessio Ansuini, Alessandro Laio, Jakob H. Macke, and Davide Zoccolan. 2019. https://arxiv.org/abs/1905.12784 Intrinsic dimension of data representations in deep neural networks . In Neural Information Processing Systems (NeurIPS)

Show all 85 references
  1. [9]

    Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2017. https://openreview.net/forum?id=SyK00v5xx A Simple but Tough-to-Beat Baseline for Sentence Embeddings . In International Conference on Learning Representations (ICLR)

  2. [10]

    Akari Asai, Timo Schick, Patrick Lewis, Xilun Chen, Gautier Izacard, Sebastian Riedel, Hannaneh Hajishirzi, and Wen-tau Yih. 2023. https://doi.org/10.18653/v1/2023.findings-acl.225 Task-aware Retrieval with Instructions . In Findings of the Association for Computational Lingui...

  3. [11]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. https://openreview.net/forum?id=IW1PR7vEBf LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders . In First Conference on Language Modeling (COLM)

  4. [12]

    Bowman, Gabor Angeli, Christopher Potts, and Christopher D

    Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. https://doi.org/10.18653/v1/D15-1075 A large annotated corpus for learning natural language inference . In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processin...

  5. [13]

    Jörg Bruske and Gerald Sommer. 1998. https://doi.org/10.1109/34.682189 Intrinsic dimensionality estimation with optimally topology preserving maps . IEEE Transactions on Pattern Analysis and Machine Intelligence, 20(5):572--575

  6. [14]

    Daniel Cer, Mona Diab, Eneko Agirre, I \ n igo Lopez-Gazpio, and Lucia Specia. 2017. https://doi.org/10.18653/v1/S17-2001 SemEval-2017 Task 1: Semantic Textual Similarity Multilingual and Crosslingual Focused Evaluation . In Proceedings of the 11th International Workshop on Se...

  7. [15]

    Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised Cross-lingual Representation Learning ...

  8. [16]

    Alexis Conneau, Douwe Kiela, Holger Schwenk, Lo \" c Barrault, and Antoine Bordes. 2017. https://doi.org/10.18653/v1/D17-1070 Supervised Learning of Universal Sentence Representations from Natural Language Inference Data . In Proceedings of the 2017 Conference on Empirical Met...

  9. [17]

    Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge

    Gabriel de Souza P. Moreira, Radek Osmulski, Mengyao Xu, Ronay Ak, Benedikt Schifferer, and Even Oldridge. 2024. https://arxiv.org/abs/2407.15831 NV-Retriever: Improving text embedding models with effective hard-negative mining . arXiv:2407.15831

  10. [18]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associati...

  11. [19]

    Georgiana Dinu, Corey Barrett, Yi Xiang, Miguel Romero Calvo, Anna Currey, and Xing Niu. 2025. https://openreview.net/forum?id=szRmEM8Kx5 Effective post-training embedding compression via temperature control in contrastive training . In International Conference on Learning Rep...

  12. [20]

    Kawin Ethayarajh. 2018. https://doi.org/10.18653/v1/W18-3012 Unsupervised Random Walk Sentence Embeddings: A Strong but Simple Baseline . In Proceedings of the Third Workshop on Representation Learning for NLP (RepL4NLP) , pages 91--100

  13. [21]

    Elena Facco, Maria d’Errico, Alex Rodriguez, and Alessandro Laio. 2017. https://arxiv.org/abs/1803.06992 Estimating the intrinsic dimension of datasets by a minimal neighborhood information . Scientific Reports, 7

  14. [22]

    Keinosuke Fukunaga and David R. Olsen. 1971. https://doi.org/10.1109/T-C.1971.223208 An Algorithm for Finding Intrinsic Dimensionality of Data . IEEE Transactions on Computers, C-20(2):176--183

  15. [23]

    Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. https://aclanthology.org/2021.emnlp-main.552 SimCSE: Simple Contrastive Learning of Sentence Embeddings . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6894--6910

  16. [24]

    Gregor Geigle, Nils Reimers, Andreas Rücklé, and Iryna Gurevych. 2021. https://arxiv.org/abs/2104.07081 TWEAC: Transformer with Extendable QA Agent Classifiers . arxiv:2104.07081, arXiv:2104.07081

  17. [25]

    Faegheh Hasibi, Fedor Nikolaev, Chenyan Xiong, Krisztian Balog, Svein Erik Bratsberg, Alexander Kotov, and Jamie Callan. 2017. https://doi.org/10.1145/3077136.3080751 DBpedia-Entity V2: A Test Collection for Entity Search . In Proceedings of the 40th International ACM SIGIR Co...

  18. [26]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. https://openreview.net/forum?id=nZeVKeeFYf9 LoRA: Low-Rank Adaptation of Large Language Models . In International Conference on Learning Representations (ICLR)

  19. [27]

    Junjie Huang, Duyu Tang, Wanjun Zhong, Shuai Lu, Linjun Shou, Ming Gong, Daxin Jiang, and Nan Duan. 2021. https://doi.org/10.18653/v1/2021.findings-emnlp.23 WhiteningBERT: An Easy Unsupervised Sentence Embedding Approach . In Findings of the Association for Computational Lingu...

  20. [28]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  21. [29]

    Ting Jiang, Shaohan Huang, Zhongzhi Luan, Deqing Wang, and Fuzhen Zhuang. 2024. https://aclanthology.org/2024.findings-emnlp.181 Scaling Sentence Embeddings with Large Language Models . In Findings of the Association for Computational Linguistics: EMNLP 2024

  22. [30]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://aclanthology.org/2020.emnlp-main.550 Dense Passage Retrieval for Open-Domain Question Answering . In Proceedings of the 2020 Conference on Empirical ...

  23. [31]

    Phillip Keung, Yichao Lu, Gy \"o rgy Szarvas, and Noah A. Smith. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.369 The Multilingual Amazon Reviews Corpus . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 4563--4568

  24. [32]

    Jaeyoung Kim, Dohyeon Lee, and Seung-won Hwang. 2024. https://doi.org/10.18653/v1/2024.naacl-long.437 HIL: Hybrid Isotropy Learning for Zero-shot Performance in Dense retrieval . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa...

  25. [33]

    Kakade, Prateek Jain, and Ali Farhadi

    Aditya Kusupati, Gantavya Bhatt, Aniket Rege, Matthew Wallingford, Aditya Sinha, Vivek Ramanujan, William Howard-Snyder, Kaifeng Chen, Sham M. Kakade, Prateek Jain, and Ali Farhadi. 2022. https://arxiv.org/abs/2205.13147 Matryoshka Representation Learning . In Proceedings of t...

  26. [34]

    Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav...

  27. [35]

    Chankyu Lee, Rajarshi Roy, Mengyao Xu, Jonathan Raiman, Mohammad Shoeybi, Bryan Catanzaro, and Wei Ping. 2024 a . https://arxiv.org/abs/2405.17428 NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models . arXiv:2405.17428

  28. [36]

    Jinhyuk Lee, Zhuyun Dai, Xiaoqi Ren, Blair Chen, Daniel Cer, Jeremy R. Cole, Kai Hui, Michael Boratko, Rajvi Kapadia, Wen Ding, Yi Luan, Sai Meher Karthik Duddu, Gustavo Hernandez Abrego, Weiqiang Shi, Nithi Gupta, Aditya Kusupati, Prateek Jain, Siddhartha Reddy Jonnalagadda, ...

  29. [37]

    Yibin Lei, Di Wu, Tianyi Zhou, Tao Shen, Yu Cao, Chongyang Tao, and Andrew Yates. 2024. https://doi.org/10.18653/v1/2024.acl-long.546 Meta-Task Prompting Elicits Embeddings from Large Language Models . In Proceedings of the 62nd Annual Meeting of the Association for Computatio...

  30. [38]

    Elizaveta Levina and Peter Bickel. 2004. https://proceedings.neurips.cc/paper_files/paper/2004/file/74934548253bcab8490ebd74afed7031-Paper.pdf Maximum Likelihood Estimation of Intrinsic Dimension . In Advances in Neural Information Processing Systems (NIPS), volume 17

  31. [39]

    Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.733 On the Sentence Embeddings from Pre-trained Language Models . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing...

  32. [40]

    Xianming Li, Zongxi Li, Jing Li, Haoran Xie, and Qing Li. 2025. https://openreview.net/forum?id=plgLA2YBLH ESE: Espresso Sentence Embeddings . In The Thirteenth International Conference on Learning Representations (ICLR)

  33. [41]

    Zehan Li, Xin Zhang, Yanzhao Zhang, Dingkun Long, Pengjun Xie, and Meishan Zhang. 2023. https://arxiv.org/abs/2308.03281 Towards General Text Embeddings with Multi-stage Contrastive Learning . arXiv:2308.03281

  34. [42]

    Maas, Raymond E

    Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. https://aclanthology.org/P11-1015/ Learning Word Vectors for Sentiment Analysis . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: H...

  35. [43]

    Marco Marelli, Stefano Menini, Marco Baroni, Luisa Bentivogli, Raffaella Bernardi, and Roberto Zamparelli. 2014. http://www.lrec-conf.org/proceedings/lrec2014/pdf/363_Paper.pdf A SICK cure for the evaluation of compositional distributional semantic models . In Proceedings of t...

  36. [44]

    Julian McAuley and Jure Leskovec. 2013. https://dl.acm.org/doi/10.1145/2507157.2507163 Hidden factors and hidden topics: understanding rating dimensions with review text

  37. [45]

    Leland McInnes, John Healy, and James Melville. 2020. https://arxiv.org/abs/1802.03426 UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction . arXiv:1802.03426

  38. [46]

    Timothee Mickus, Stig-Arne Gr \"o nroos, and Joseph Attieh. 2024. https://doi.org/10.18653/v1/2024.acl-short.7 Isotropy, Clusters, and Classifiers . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 75--84

  39. [47]

    Jiaqi Mu and Pramod Viswanath. 2018. https://openreview.net/forum?id=HkuGJ3kCb All-but-the-Top: Simple and Effective Postprocessing for Word Representations . In International Conference on Learning Representations (ICLR)

  40. [48]

    Niklas Muennighoff. 2022. https://doi.org/10.48550/ARXIV.2202.08904 SGPT: GPT Sentence Embeddings for Semantic Search . arXiv:2202.08904

  41. [49]

    Niklas Muennighoff, Hongjin Su, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. 2024. https://arxiv.org/abs/2402.09906 Generative Representational Instruction Tuning . arXiv:2402.09906

  42. [50]

    Niklas Muennighoff, Nouamane Tazi, Loic Magne, and Nils Reimers. 2023. https://doi.org/10.18653/v1/2023.eacl-main.148 MTEB: Massive Text Embedding Benchmark . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics (EACL),...

  43. [51]

    Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. https://arxiv.org/abs/1611.09268 MS MARCO: A Human Generated MAchine Reading COmprehension Dataset . CoRR, abs/1611.09268

  44. [52]

    Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith Hall, Daniel Cer, and Yinfei Yang. 2022 a . https://doi.org/10.18653/v1/2022.findings-acl.146 Sentence-T5: Scalable Sentence Encoders from Pre-trained Text-to-Text Models . In Findings of the Association for Comp...

  45. [53]

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022 b . https://doi.org/10.18653/v1/2022.emnlp-main.669 Large Dual Encoders Are Generalizable Retrievers . In Proceedings of the 2022 ...

  46. [54]

    Morris, Brandon Duderstadt, and Andriy Mulyar

    Zach Nussbaum, John X. Morris, Brandon Duderstadt, and Andriy Mulyar. 2024. https://arxiv.org/abs/2402.01613 Nomic Embed: Training a Reproducible Long Context Text Embedder . arXiv:2402.01613

  47. [55]

    James O ' Neill, Polina Rozenshtein, Ryuichi Kiryo, Motoko Kubota, and Danushka Bollegala. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.568 I Wish I Would Have Loved This One, But I Didn`t -- A Multilingual Dataset for Counterfactual Detection in Product Review . In Proce...

  48. [56]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  49. [57]

    Nils Reimers and Iryna Gurevych. 2019. https://doi.org/10.18653/v1/D19-1410 Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on ...

  50. [58]

    Andrew Rosenberg and Julia Hirschberg. 2007. https://aclanthology.org/D07-1043 V-Measure: A Conditional Entropy-Based External Cluster Evaluation Measure . In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural...

  51. [59]

    Rousseeuw

    Peter J. Rousseeuw. 1987. https://doi.org/10.1016/0377-0427(87)90125-7 Silhouettes: A graphical aid to the interpretation and validation of cluster analysis . Journal of Computational and Applied Mathematics, 20:53--65

  52. [60]

    William Rudman and Carsten Eickhoff. 2024. https://openreview.net/forum?id=dbQH9AOVd5 Stable Anisotropic Regularization . In The Twelfth International Conference on Learning Representations (ICLR)

  53. [61]

    William Rudman, Nate Gillman, Taylor Rayne, and Carsten Eickhoff. 2022. https://doi.org/10.18653/v1/2022.findings-acl.262 IsoScore: Measuring the Uniformity of Embedding Space Utilization . In Findings of the Association for Computational Linguistics: ACL 2022, pages 3325--3339

  54. [62]

    Dinghan Shen, Guoyin Wang, Wenlin Wang, Martin Renqiang Min, Qinliang Su, Yizhe Zhang, Chunyuan Li, Ricardo Henao, and Lawrence Carin. 2018. https://doi.org/10.18653/v1/P18-1041 Baseline Needs More Love: On Simple Word-Embedding-Based Models and Associated Pooling Mechanisms ....

  55. [63]

    Jacob Mitchell Springer, Suhas Kotha, Daniel Fried, Graham Neubig, and Aditi Raghunathan. 2024. https://arxiv.org/abs/2402.15449 Repetition Improves Language Model Embeddings . arXiv:2402.15449

  56. [64]

    Smith, Luke Zettlemoyer, and Tao Yu

    Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu. 2023. https://doi.org/10.18653/v1/2023.findings-acl.71 One Embedder, Any Task: Instruction-Finetuned Text Embeddings . In Findings of the Asso...

  57. [65]

    Jianlin Su, Jiarun Cao, Weijie Liu, and Yangyiwen Ou. 2021. https://arxiv.org/abs/2103.15316 Whitening Sentence Representations for Better Semantics and Faster Retrieval . arXiv:2103.15316

  58. [66]

    Tenenbaum, Vin de Silva, and John C

    Joshua B. Tenenbaum, Vin de Silva, and John C. Langford. 2000. https://doi.org/10.1126/science.290.5500.2319 A Global Geometric Framework for Nonlinear Dimensionality Reduction . Science, 290(5500):2319--2323

  59. [67]

    Hayato Tsukagoshi, Ryohei Sasano, and Koichi Takeda. 2021. https://doi.org/10.18653/v1/2021.acl-short.52 DefSent: Sentence Embeddings using Definition Sentences . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Internatio...

  60. [68]

    Eduard Tulchinskii, Kristian Kuznetsov, Kushnareva Laida, Daniil Cherniavskii, Sergey Nikolenko, Evgeny Burnaev, Serguei Barannikov, and Irina Piontkovskaya. 2023. https://openreview.net/forum?id=8uOZ0kNji6 Intrinsic Dimension Estimation for Robust Detection of AI -Generated T...

  61. [69]

    Laurens van der Maaten and Geoffrey E. Hinton. 2008. https://www.jmlr.org/papers/volume9/vandermaaten08a/vandermaaten08a.pdf Visualizing Data using t-SNE . Journal of Machine Learning Research, 9:2579--2605

  62. [70]

    Feng Wang and Huaping Liu. 2021. https://openaccess.thecvf.com/content/CVPR2021/papers/Wang_Understanding_the_Behaviour_of_Contrastive_Loss_CVPR_2021_paper.pdf Understanding the Behaviour of Contrastive Loss . pages 2495--2504

  63. [71]

    Hongwei Wang, Hongming Zhang, and Dong Yu. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.694 On the Dimensionality of Sentence Embeddings . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 10344--10354

  64. [72]

    Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. https://arxiv.org/abs/2212.03533 Text Embeddings by Weakly-Supervised Contrastive Pre-training . arXiv:2212.03533

  65. [73]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 a . https://doi.org/10.18653/v1/2024.acl-long.642 Improving Text Embeddings with Large Language Models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Lingui...

  66. [74]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024 b . https://arxiv.org/abs/2402.05672 Multilingual e5 text embeddings: A technical report . arXiv:2402.05672

  67. [75]

    Adina Williams, Nikita Nangia, and Samuel Bowman. 2018. https://doi.org/10.18653/v1/N18-1101 A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference . In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computationa...

  68. [76]

    Chenghao Xiao, Yang Long, and Noura Al Moubayed. 2023. https://doi.org/10.18653/v1/2023.findings-acl.778 On Isotropy, Contextualization and Learning Dynamics of Contrastive-based Sentence Representation Learning . In Findings of the Association for Computational Linguistics: A...

  69. [77]

    Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2024. https://arxiv.org/abs/2309.07597 C-Pack: Packaged Resources To Advance General Chinese Embedding . In The 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), ...

  70. [78]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, and 43 others. 2024. https://arxiv.org/ab...

  71. [79]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. https://doi.org/10.18653/v1/D18-1259 HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering . In Proceedings of the 2018 Conference on...

  72. [80]

    Chihiro Yano, Akihiko Fukuchi, Shoko Fukasawa, Hideyuki Tachibana, and Yotaro Watanabe. 2024. https://aclanthology.org/2024.lrec-main.1034 Multilingual Sentence-T5: Scalable Sentence Encoders for Multilingual Applications . In Proceedings of the 2024 Joint International Confer...

  73. [81]

    Sho Yokoi, Han Bao, Hiroto Kurita, and Hidetoshi Shimodaira. 2024. https://proceedings.neurips.cc/paper_files/paper/2024/file/dd1fef536655685898a6602bfbf16857-Paper-Conference.pdf Zipfian whitening . In Advances in Neural Information Processing Systems (NeurIPS), volume 37, pa...

  74. [82]

    Xinyu Zhang, Nandan Thakur, Odunayo Ogundepo, Ehsan Kamalloo, David Alfonso-Hermelo, Xiaoguang Li, Qun Liu, Mehdi Rezagholizadeh, and Jimmy Lin. 2022. https://arxiv.org/abs/2210.09984 Making a MIRACL: Multilingual Information Retrieval Across a Continuum of Languages . arXiv:2...

  75. [83]

    Wenjie Zhuo, Yifan Sun, Xiaohan Wang, Linchao Zhu, and Yi Yang. 2023. https://doi.org/10.18653/v1/2023.acl-long.677 WhitenedCSE: Whitening-based Contrastive Learning of Sentence Embeddings . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguis...

  76. [84]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  77. [85]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.