{"work":{"id":"27adfcc9-2a67-43d6-a844-78309012411f","openalex_id":"https://openalex.org/W2971193649","doi":"10.48550/arxiv.1908.10084","arxiv_id":"1908.10084","raw_key":null,"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","authors":null,"authors_text":"Nils Reimers, Iryna Gurevych","year":2019,"venue":"cs.CL","abstract":"BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering.\n  In this publication, we present Sentence-BERT (SBERT), a modification of the pretrained BERT network that use siamese and triplet network structures to derive semantically meaningful sentence embeddings that can be compared using cosine-similarity. This reduces the effort for finding the most similar pair from 65 hours with BERT / RoBERTa to about 5 seconds with SBERT, while maintaining the accuracy from BERT.\n  We evaluate SBERT and SRoBERTa on common STS tasks and transfer learning tasks, where it outperforms other state-of-the-art sentence embeddings methods.","external_url":"https://arxiv.org/abs/1908.10084","cited_by_count":88,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1908.10084","created_at":"2026-05-08T18:39:01.827291+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","render_title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks"},"hub":{"state":{"work_id":"27adfcc9-2a67-43d6-a844-78309012411f","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":280,"external_cited_by_count":88,"distinct_field_count":26,"first_pith_cited_at":"2019-10-09T03:23:22+00:00","last_pith_cited_at":"2026-07-09T00:27:07+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T00:09:30.724713+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":17},{"context_role":"method","n":16},{"context_role":"other","n":2},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"use_method","n":16},{"context_polarity":"background","n":13},{"context_polarity":"unclear","n":4},{"context_polarity":"support","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","claims":[{"claim_text":"BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering.\n  In this publication, we present Sentence-BER","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:00:02.494205+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"f3fe8546-e2f0-4190-889c-f65e4446dd58","orcid":null,"display_name":"Nils Reimers"},{"id":"39c270ca-fad8-48d4-a82a-63645919b5d5","orcid":null,"display_name":"Iryna Gurevych"}]},"error":null,"updated_at":"2026-05-14T17:59:27.741459+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T05:26:30.875351+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":13},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":13},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":9},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":8},{"title":"Simcse: Simple contrastive learning of sentence embeddings","work_id":"e9fab1e4-f443-4963-9f2a-83f772482c00","shared_citers":8},{"title":"Text Embeddings by Weakly-Supervised Contrastive Pre-training","work_id":"789cc674-467e-4f23-bb50-05c79fe8c4c2","shared_citers":8},{"title":"Passage Re-ranking with BERT","work_id":"562fbfab-d6fe-48e1-a06d-e5d078c70945","shared_citers":7},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":6},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","work_id":"bab684a8-d933-426c-a19e-2c855a0d1f59","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"Retrieval-Augmented Generation for Large Language Models: A Survey","work_id":"b80d2790-6cd9-4c87-b3c4-de404f99a80e","shared_citers":6},{"title":"arXiv preprint arXiv:2004.04906 , year=","work_id":"3d6f2008-b001-4542-ba3f-192f6880c74b","shared_citers":5},{"title":"BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension","work_id":"7ab72623-0d1b-41fc-96e9-181246f2ea00","shared_citers":5},{"title":"BERTScore: Evaluating Text Generation with BERT","work_id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","shared_citers":5},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":5},{"title":"Finetuned Language Models Are Zero-Shot Learners","work_id":"7ed6cdaa-ed67-4db4-aceb-b7e1b0e6e7c4","shared_citers":5},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":5},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"Mteb: Massive text embedding benchmark","work_id":"7423c7ba-a393-45d6-bb35-9bc2c931a545","shared_citers":5}],"time_series":[{"n":1,"year":2019},{"n":1,"year":2021},{"n":2,"year":2022},{"n":2,"year":2023},{"n":2,"year":2024},{"n":97,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T05:26:27.489946+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T05:26:48.623076+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","claims":[{"claim_text":"BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering.\n  In this publication, we present Sentence-BER","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T18:00:06.520745+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","claims":[{"claim_text":"BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering.\n  In this publication, we present Sentence-BER","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T05:26:26.745415+00:00"}},"summary":{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","claims":[{"claim_text":"BERT (Devlin et al., 2018) and RoBERTa (Liu et al., 2019) has set a new state-of-the-art performance on sentence-pair regression tasks like semantic textual similarity (STS). However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT. The construction of BERT makes it unsuitable for semantic similarity search as well as for unsupervised tasks like clustering.\n  In this publication, we present Sentence-BER","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":13},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":13},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":9},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":8},{"title":"Simcse: Simple contrastive learning of sentence embeddings","work_id":"e9fab1e4-f443-4963-9f2a-83f772482c00","shared_citers":8},{"title":"Text Embeddings by Weakly-Supervised Contrastive Pre-training","work_id":"789cc674-467e-4f23-bb50-05c79fe8c4c2","shared_citers":8},{"title":"Passage Re-ranking with BERT","work_id":"562fbfab-d6fe-48e1-a06d-e5d078c70945","shared_citers":7},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":6},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","work_id":"bab684a8-d933-426c-a19e-2c855a0d1f59","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"Retrieval-Augmented Generation for Large Language Models: A Survey","work_id":"b80d2790-6cd9-4c87-b3c4-de404f99a80e","shared_citers":6},{"title":"arXiv preprint arXiv:2004.04906 , year=","work_id":"3d6f2008-b001-4542-ba3f-192f6880c74b","shared_citers":5},{"title":"BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension","work_id":"7ab72623-0d1b-41fc-96e9-181246f2ea00","shared_citers":5},{"title":"BERTScore: Evaluating Text Generation with BERT","work_id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","shared_citers":5},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":5},{"title":"Finetuned Language Models Are Zero-Shot Learners","work_id":"7ed6cdaa-ed67-4db4-aceb-b7e1b0e6e7c4","shared_citers":5},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":5},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"Mteb: Massive text embedding benchmark","work_id":"7423c7ba-a393-45d6-bb35-9bc2c931a545","shared_citers":5}],"time_series":[{"n":1,"year":2019},{"n":1,"year":2021},{"n":2,"year":2022},{"n":2,"year":2023},{"n":2,"year":2024},{"n":97,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"39c270ca-fad8-48d4-a82a-63645919b5d5","orcid":null,"display_name":"Iryna Gurevych","source":"manual","import_confidence":0.72},{"id":"f3fe8546-e2f0-4190-889c-f65e4446dd58","orcid":null,"display_name":"Nils Reimers","source":"manual","import_confidence":0.72}]}}