Pith. sign in

REVIEW 19 cited by

Whitening Sentence Representations for Better Semantics and Faster Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.15316 v1 pith:GHP3J7IE submitted 2021-03-29 cs.CL cs.AIcs.LG

Whitening Sentence Representations for Better Semantics and Faster Retrieval

classification cs.CL cs.AIcs.LG
keywords sentencemodelrepresentationrepresentationswhiteningachieveachievedbetter
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Pre-training models such as BERT have achieved great success in many natural language processing tasks. However, how to obtain better sentence representation through these pre-training models is still worthy to exploit. Previous work has shown that the anisotropy problem is an critical bottleneck for BERT-based sentence representation which hinders the model to fully utilize the underlying semantic features. Therefore, some attempts of boosting the isotropy of sentence distribution, such as flow-based model, have been applied to sentence representations and achieved some improvement. In this paper, we find that the whitening operation in traditional machine learning can similarly enhance the isotropy of sentence representations and achieve competitive results. Furthermore, the whitening technique is also capable of reducing the dimensionality of the sentence representation. Our experimental results show that it can not only achieve promising performance but also significantly reduce the storage cost and accelerate the model retrieval speed.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SimCSE: Simple Contrastive Learning of Sentence Embeddings

    cs.CL 2021-04 conditional novelty 8.0

    SimCSE achieves 76.3% unsupervised and 81.6% supervised Spearman's correlation on STS tasks with BERT-base, improving prior best results by 4.2% and 2.2% via simple contrastive learning.

  2. Anisotropy Decides Cosine vs. Rank Metrics for Text Embeddings

    cs.CL 2026-06 conditional novelty 7.0

    Anisotropy, quantified by dominant-dimension variance fraction, determines the best parameter-free similarity metric for text embeddings, with rank-based metrics gaining ~20% relative where cosine is weakest.

  3. CrossAlpha: An Annual-Report Benchmark for Cross-Market Factor Researc (with LLM Agents)

    cs.IR 2026-05 unverdicted novelty 7.0

    CrossAlpha is a new open benchmark covering 3600 firms across US, Japan, Taiwan, South Korea and Hong Kong that distills reports into 10-category descriptions, constructs residual cross-market similarity graphs, and t...

  4. Concepts Whisper While Syntax Shouts: Spectral Anti-Concentration and the Dual Geometry of Transformer Representations

    cs.LG 2026-05 unverdicted novelty 7.0

    Transformer activations show spectral anti-concentration for concepts in the tail while syntax prefers high-variance directions, forming a dual geometry.

  5. Spectral Tempering for Embedding Compression in Dense Passage Retrieval

    cs.IR 2026-03 unverdicted novelty 7.0

    Spectral Tempering derives an adaptive scaling factor γ(k) from the embedding eigenspectrum via local SNR analysis and knee-point normalization to achieve near-optimal compression without training or validation.

  6. Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis

    stat.ME 2026-07 accept novelty 6.0

    A scale-free gain rule applied to variational ELBO paths recovers true factor dimensionality in partially exploratory factor analysis where raw information criteria over-factor.

  7. Semantic Code Clone Detection: Are We There Yet?

    cs.SE 2026-06 conditional novelty 6.0

    SOTA semantic code clone detectors exhibit substantial performance degradation on distribution-shifted yet semantically equivalent clones, revealing reliance on lexical and structural shortcuts rather than semantic un...

  8. When Global Gating Is Enough: Admission-Time Hubness Control in Anisotropic Vector Retrieval Systems

    cs.CR 2026-06 unverdicted novelty 6.0

    Global admission-time gating with sentinel queries controls vector hubness in anisotropic embeddings, achieving recall 1.0 on critical attack points and 0.91 on HotFlip attacks with 1% false positives, while per-topic...

  9. Semantic Identification of IoT Devices from Behavioral Primitives

    cs.CR 2026-06 unverdicted novelty 6.0

    Semantic matching of MUD ACE behavioral primitives yields more robust IoT device identification than exact overlap under runtime variations and sparse observations.

  10. Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings

    cs.CL 2026-06 unverdicted novelty 6.0

    EmbedFilter applies a linear filter derived from the LLM unembedding matrix to suppress high-frequency token influences in text embeddings, yielding improved zero-shot performance and inherent dimensionality reduction.

  11. Context Memorization for Efficient Long Context Generation

    cs.CL 2026-05 unverdicted novelty 6.0

    Attention-state memory externalizes long prefixes into a lightweight lookup table of precomputed attention states, yielding higher accuracy than standard in-context learning at fixed memory budgets and lower latency t...

  12. ASPIRE: Make Spectral Graph Collaborative Filtering Great Again via Adaptive Filter Learning

    cs.IR 2026-04 unverdicted novelty 6.0

    ASPIRE learns adaptive graph filters via bi-level optimization to overcome low-frequency explosion bias in spectral collaborative filtering, achieving strong performance and stability.

  13. REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

    cs.CL 2026-04 unverdicted novelty 6.0

    REZE controls representation shifts in contrastive pre-finetuning of text embeddings via eigenspace decomposition of anchor-positive pairs and adaptive soft-shrinkage on task-variant directions.

  14. Text and Code Embeddings by Contrastive Pre-Training

    cs.CL 2022-01 unverdicted novelty 6.0

    Contrastive pre-training on unsupervised data at scale creates text and code embeddings that set new state-of-the-art results on classification and semantic search benchmarks.

  15. Correlation Is Not Enough: Embedding Human Metadata for Individual Causal Discovery

    cs.AI 2026-06 unverdicted novelty 5.0

    Contrastive fine-tuning on 72k pairs plus BODHI hard-negative mining from a biomedical KG improves PubMedBERT separation from 1.05x to 2.30x and cross-domain discrimination by +0.392 while preserving most BIOSSES performance.

  16. Towards General Text Embeddings with Multi-stage Contrastive Learning

    cs.CL 2023-08 unverdicted novelty 5.0

    GTE_base is a compact text embedding model using multi-stage contrastive learning on diverse data that outperforms OpenAI's API and 10x larger models on massive benchmarks and works for code as text.

  17. MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation

    cs.CL 2026-06 unverdicted novelty 3.0

    MODE-RAG introduces a VFE-driven multi-agent pipeline with MCTS and logit perturbations to lower hallucination and sycophancy rates in multimodal RAG, tested on the new ModeVent subset of MultiVent.

  18. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 2.0

    A comprehensive reference book organizing existing techniques for agentic AI systems across LLM substrate, reasoning, agent design patterns, inter-agent coordination, and production deployment.

  19. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

    cs.AI 2026-06 unverdicted novelty 1.0

    A survey-style reference book mapping the full agentic-AI stack from transformer internals to production deployment, with no new research result.