Pith. sign in

REVIEW 37 cited by

SGPT: GPT Sentence Embeddings for Semantic Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.08904 v5 pith:BS67I6TO submitted 2022-02-17 cs.CL cs.AIcs.IR

classification cs.CLcs.AIcs.IR
keywords embeddingssearchsentencesgptmodelsparameterssemanticbillion
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Decoder transformers have continued increasing in scale reaching hundreds of billions of parameters. Due to their scale the same decoder sets state-of-the-art results on various language tasks via prompting or fine-tuning. Yet, these large foundation models remain unusable for the related fields of semantic search and sentence embeddings. This prevents possibly new state-of-the-art results and forces organizations to train and maintain separate models. To this end, we propose SGPT to use decoders for sentence embeddings and semantic search via prompting or fine-tuning. At 5.8 billion parameters SGPT improves on the previously best sentence embeddings by a margin of 7% and outperforms a concurrent method with 175 billion parameters as measured on the BEIR search benchmark. Code, models and result files are freely available at https://github.com/Muennighoff/sgpt.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 37 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

    cs.CL 2026-08 conditional novelty 7.0 of 10

    Tevatron-Elastic unifies depth, token, and width compression for retrievers and rerankers into one abstraction that reproduces prior elastic methods as special cases and adds a new multi-ratio token compression method (MLTC).

  2. Leveraging Large Vision-Language Model as User Intent-aware Encoder for Composed Image Retrieval

    cs.IR 2024-12 conditional novelty 7.0 of 10

    CIR-LVLM fine-tunes Qwen-VL-Chat with LoRA and hybrid task and instance-specific prompts to produce query and target embeddings, achieving new state-of-the-art recall on Fashion-IQ, Shoes, and CIRR.

  3. Inference Scaling for Bridging Retrieval and Augmented Generation

    cs.CL 2024-12 conditional novelty 7.0 of 10

    MOI estimates a debiased utility for each retrieved passage from multiple permuted reads and reranks by it, yielding large RAG quality gains at the cost of extra LLM calls.

  4. The Embedder's Dilemma: LLMs Are Better, but at What Cost?

    cs.CL 2026-08 conditional novelty 6.0 of 10

    Across 37 tasks the best LLM and best embedding model score about the same (77.6 vs 77.2), but the LLM costs roughly 1,400 times more, with LLMs winning only on reasoning-heavy retrieval.

  5. Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment $\unicode{x2013}$ Is English Enough?

    cs.CL 2026-08 conditional novelty 6.0 of 10

    English-based cross-lingual alignment predicts LLM translation quality as well as or better than direct source-target alignment, supporting the English-pivot hypothesis.

  6. IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment

    cs.CL 2026-07 conditional novelty 6.0 of 10

    IRIS extracts identity embeddings from frozen LLMs so each entity is encoded once from its own knowledge graph and matched to other graphs by cosine similarity, hitting 97.99-100.00 Hits@1 on four benchmarks.

  7. BitNet Text Embeddings

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    BITEMBED trains 1.58-bit ternary-weight LLM embedders with contrastive pre-training, supervised distillation, and multi-precision output training, matching FP16 teachers within ~0.6 MMTEB points at ~2x CPU speed.

  8. BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels

    cs.IR 2026-04 unverdicted novelty 6.0 of 10

    BioHiCL applies hierarchical multi-label contrastive learning with MeSH annotations to improve biomedical retrieval, sentence similarity, and question answering using small efficient models.

  9. A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't)

    cs.LG 2026-02 conditional novelty 6.0 of 10

    Only gradient-based (LESS) representations make subset-to-query distance a reliable predictor of instruction-tuning performance; greedy round-robin helps most at small budgets, and random selection is surprisingly com...

  10. LLM-based Embeddings: Attention Values Encode Sentence Semantics Better Than Hidden States

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Pooling attention value vectors gives better training-free LLM sentence embeddings than pooling hidden states.

  11. LAMDAS: LLM as an Implicit Classifier for Domain-specific Data Selection

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LAMDAS selects domain-relevant training data via an LLM likelihood ratio with a learned domain prefix, beating full-data training and nine baselines on code and math.

  12. Delta Activations: A Representation for Finetuned Large Language Models

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Delta Activations embed finetuned LLMs as the average difference in hidden states between the finetuned model and its base model on a small set of generic prompts, yielding domain clusters and approximate additive com...

  13. Negative Matters: Multi-Granularity Hard-Negative Synthesis and Anchor-Token-Aware Pooling for Enhanced Text Embeddings

    cs.CL 2025-08 conditional novelty 6.0 of 10

    A new MTEB state-of-the-art for text embeddings is reported by combining multi-granularity LLM-generated hard negatives with curriculum training and an anchor-token-aware pooling method.

  14. CRED-SQL: Enhancing Real-world Large Scale Database Text-to-SQL Parsing through Cluster Retrieval and Execution Description

    cs.CL 2025-08 conditional novelty 6.0 of 10

    CRED-SQL substantially improves Text-to-SQL on large-schema benchmarks by down-weighting common schema columns during retrieval and generating SQL through a natural language execution description.

  15. LLM2Rec: Large Language Models Are Powerful Embedding Models for Sequential Recommendation

    cs.IR 2025-06 conditional novelty 6.0 of 10

    LLM2Rec combines next-item prediction fine-tuning with masked token reconstruction and contrastive learning to produce item embeddings that outperform existing text-embedding baselines for sequential recommendation.

  16. Maximally-Informative Retrieval for State Space Model Generation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    RICO ranks documents by how much they reduce an SSM's question perplexity, using gradient-document inner products, and matches BM25 while often beating E5 on answer quality without finetuning.

  17. Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Prompt-based text embeddings can be truncated to a small fraction of their dimensions with little performance loss on classification and clustering, but retrieval and STS degrade faster; the difference tracks lower in...

  18. How Programming Concepts and Neurons Are Shared in Code Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    In Llama-based code models, programming languages are represented through an English-like intermediate token space, with language-specific neurons concentrated in bottom layers and exclusive PL neurons in top layers; ...

  19. DeepRTL2: A Versatile Model for RTL-Related Tasks

    cs.AR 2025-05 reject novelty 6.0 of 10

    DeepRTL2 claims state-of-the-art results across RTL generation, understanding, code search, equivalence checking, and performance prediction, but the evidence is weakened by benchmark construction issues and a contrad...

  20. Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Large language models, especially GPT-4 with few-shot prompts, can classify topological spatial relations between WKT-encoded geometries with roughly 0.6 to 0.66 accuracy, though errors cluster near conceptually simil...

  21. Memorization and Knowledge Injection in Gated LLMs

    cs.CL 2025-04 conditional novelty 6.0 of 10

    MEGa injects episodic memories into separate gated LoRA adapters selected by embedding similarity, mitigating catastrophic forgetting and enabling recall, QA, and compositional questions on two datasets.

  22. Large Language Model Can Be a Foundation for Hidden Rationale-Based Retrieval

    cs.IR 2024-12 conditional novelty 6.0 of 10

    A cross-encoder LLM prompted with a binary relevance question and scored by next-token probabilities outperforms similarity-based retrievers on hidden rationale retrieval tasks.

  23. Towards modeling evolving longitudinal health trajectories with a transformer-based deep learning model

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A causal transformer that predicts each patient's future disease diagnoses repeatedly as their health record grows, producing a continuous risk trajectory over time.

  24. Revealing the Numeracy Gap: An Empirical Investigation of Text Embedding Models

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Across 13 embedding models and 18 numeric formats, retrieval accuracy on the new EmbedNum-1K benchmark averages 54%, just above chance, showing that embedding models largely fail to encode numeric detail.

  25. Position: Text Embeddings Should Capture Implicit Semantics, Not Just Surface Meaning

    cs.CL 2025-06 conditional novelty 5.0 of 10

    State-of-the-art text embeddings lag far behind on tasks requiring pragmatic inference, stance detection, and social meaning, relative to their strong performance on surface semantic benchmarks.

  26. Accelerating Adaptive Retrieval Augmented Generation via Instruction-Driven Representation Reduction of Retrieval Overlaps

    cs.AI 2025-05 conditional novelty 5.0 of 10

    An acceleration method for adaptive RAG that reuses cached key-value representations of overlapping documents and uses document-derived drafts for parallel decoding, achieving about 2x end-to-end speedup.

  27. ASRank: Zero-Shot Re-Ranking with Answer Scent for Document Retrieval

    cs.CL 2025-01 conditional novelty 5.0 of 10

    ASRank re-ranks retrieved documents by scoring how well each document supports a zero-shot answer scent generated by a large LLM, beating UPR and RankGPT on several QA datasets.

  28. Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning

    cs.CV 2024-11 conditional novelty 5.0 of 10

    TRIDENT improves compositional zero-shot recognition by using LLaVA hidden states as word embeddings and smoothing attribute labels with auxiliary adjectives generated by GPT-3.5.

  29. A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement

    cs.CL 2026-08 conditional novelty 4.0 of 10

    A Word2Vec-GMM-Random Forest pipeline detects accusatory procurement comments in Ecuador's SOCE data with 0.84 precision and 0.91 recall, but those metrics are conditional on a label-selected cluster filter.

  30. Training-Free versus Training-Based Intent Classification in LLMs: Accuracy, Robustness, and Failure Modes

    cs.CL 2026-08 conditional novelty 4.0 of 10

    Statistical classifiers built on LLM activation norms and coordinates match or beat trained MLP heads on coarse intent routing and resist camouflage better, while MLPs win on fine-grained subfield distinctions.

  31. A Multi-Task Evaluation of LLMs' Processing of Academic Text Input

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    The abstract reports Gemini underperforms on four academic text tasks, but the attached full text is an unrelated biomedical retrieval paper, leaving the claims unverifiable.

  32. From Neurons to Semantics: Evaluating Cross-Linguistic Alignment Capabilities of Large Language Models via Neurons Alignment

    cs.CL 2025-07 conditional novelty 4.0 of 10

    A neuron-activation-based alignment score for LLMs correlates highly with downstream multilingual performance and transferability across nine open models.

  33. Framework of Voting Prediction of Parliament Members

    cs.SI 2025-05 reject novelty 4.0 of 10

    A multi-country framework predicts individual parliamentary votes with up to 85% accuracy and bill outcomes with up to 84% accuracy, but the evaluation does not include trivial baselines.

  34. Jasper and Stella: distillation of SOTA embedding models

    cs.IR 2024-12 conditional novelty 4.0 of 10

    A 2B-parameter embedding model distilled from two larger teachers achieves a 71.54 average MTEB score (No.3 as of Dec 2024), matching 7B-parameter models.

  35. DynRank: Improving Passage Retrieval with Dynamic Zero-Shot Prompting Based on Question Classification

    cs.CL 2024-11 conditional novelty 4.0 of 10

    DynRank conditions UPR-style passage reranking on an automatically inferred fine-grained question type and reports small gains over static prompting on NQ, TriviaQA, WebQuestions, and BEIR.

  36. Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework

    cs.IR 2026-07 conditional novelty 3.0 of 10

    On four BEIR subsets T3EM leads nDCG@10 (0.638) but mE5-L is the recommended open default; training objective and chunk size dominate size, and no model wins every MTEB task.

  37. LLMs are Also Effective Embedding Models: An In-depth Overview

    cs.CL 2024-12 conditional novelty 2.0 of 10

    A structured survey of using decoder-only LLMs as text embedding models, covering prompting, fine-tuning, data construction, benchmarks, and open problems.

Pith tools