{"work":{"id":"bab684a8-d933-426c-a19e-2c855a0d1f59","openalex_id":"https://openalex.org/W4415354699","doi":"10.1016/j.displa.2025.103255","arxiv_id":"2506.05176","raw_key":null,"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","authors":null,"authors_text":"Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang","year":2025,"venue":"cs.CL","abstract":"In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the Qwen3 LLMs serve not only as backbone models but also play a crucial role in synthesizing high-quality, rich, and diverse training data across multiple domains and languages, thus enhancing the training pipeline. The Qwen3 Embedding series offers a spectrum of model sizes (0.6B, 4B, 8B) for both embedding and reranking tasks, addressing diverse deployment scenarios where users can optimize for either efficiency or effectiveness. Empirical evaluations demonstrate that the Qwen3 Embedding series achieves state-of-the-art results across diverse benchmarks. Notably, it excels on the multilingual evaluation benchmark MTEB for text embedding, as well as in various retrieval tasks, including code retrieval, cross-lingual retrieval and multilingual retrieval. To facilitate reproducibility and promote community-driven research and development, the Qwen3 Embedding models are publicly available under the Apache 2.0 license.","external_url":"https://arxiv.org/abs/2506.05176","cited_by_count":5,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2506.05176","created_at":"2026-05-08T17:28:41.954338+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","render_title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models"},"hub":{"state":{"work_id":"bab684a8-d933-426c-a19e-2c855a0d1f59","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":329,"external_cited_by_count":5,"distinct_field_count":22,"first_pith_cited_at":"2025-07-01T17:45:48+00:00","last_pith_cited_at":"2026-07-08T13:19:52+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T08:29:27.367033+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":14},{"context_role":"method","n":6},{"context_role":"baseline","n":3},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":11},{"context_polarity":"use_method","n":6},{"context_polarity":"baseline","n":3},{"context_polarity":"unclear","n":2},{"context_polarity":"support","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","claims":[{"claim_text":"In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T02:04:04.242568+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"8cfd545f-d741-4ca1-89e7-cc2394bfeca3","orcid":null,"display_name":"Yanzhao Zhang"},{"id":"2ddce3fd-0f51-461c-8c9a-2b3ba42d0bd1","orcid":null,"display_name":"Mingxin Li"},{"id":"6b268705-395a-45a3-8f13-05e372ec3623","orcid":null,"display_name":"Dingkun Long"},{"id":"6314c38d-420c-49be-9eff-c70b54f3eb8e","orcid":null,"display_name":"Xin Zhang"},{"id":"496bb56e-3efb-4d5a-936f-7267383c13cc","orcid":null,"display_name":"Huan Lin"},{"id":"3e0735af-daed-4e22-b48a-b7778e5e9a45","orcid":null,"display_name":"Baosong Yang"}]},"error":null,"updated_at":"2026-05-14T02:04:04.237791+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T01:54:14.613527+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":24},{"title":"Text Embeddings by Weakly-Supervised Contrastive Pre-training","work_id":"789cc674-467e-4f23-bb50-05c79fe8c4c2","shared_citers":17},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":15},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":12},{"title":"Embeddinggemma: Powerful and lightweight text representations","work_id":"e83995d5-78dc-4a7b-8d9a-1776d004351d","shared_citers":11},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":10},{"title":"Multilingual E5 Text Embeddings: A Technical Report","work_id":"b2004e20-1b24-4c37-9fad-dfdd9a7a8fee","shared_citers":9},{"title":"arXiv preprint arXiv:2312.02724 , year=","work_id":"0d9b3ad1-b405-412f-81ee-fd6f941d2367","shared_citers":8},{"title":"BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models","work_id":"c5f7f027-ac36-4b07-b824-0eca2f310641","shared_citers":8},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":8},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"Retrieval-Augmented Generation for Large Language Models: A Survey","work_id":"b80d2790-6cd9-4c87-b3c4-de404f99a80e","shared_citers":8},{"title":"doi: 10.18653/v1/N19-1423","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":7},{"title":"Gemini embedding: Generalizable embeddings from gemini","work_id":"911b6918-a128-453f-ae99-94388c38fcb1","shared_citers":7},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":7},{"title":"The probabilistic relevance framework: BM25 and beyond.Foundations and Trends in Information Retrieval, 3(4):333–389","work_id":"3dfaa21d-3751-420b-84f7-aeceda058b63","shared_citers":7},{"title":"Towards General Text Embeddings with Multi-stage Contrastive Learning","work_id":"861a61de-66fe-49d1-b1ab-11f8b082a4cc","shared_citers":7},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"Mmteb: Massive multilingual text embedding benchmark","work_id":"774aa5f1-35ad-4b36-b6cc-5f461cfab347","shared_citers":6},{"title":"MS MARCO: A Human Generated MAchine Reading COmprehension Dataset","work_id":"78d498ce-11db-4f88-8eb0-40e0f86af615","shared_citers":6},{"title":"MTEB: Massive text embedding benchmark","work_id":"5c5a0bda-f984-4ea2-8938-cab73425e1b6","shared_citers":6}],"time_series":[{"n":111,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T02:04:08.461544+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T01:54:19.373859+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","claims":[{"claim_text":"In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T02:04:08.373742+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","claims":[{"claim_text":"In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T01:54:23.950062+00:00"}},"summary":{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","claims":[{"claim_text":"In this work, we introduce the Qwen3 Embedding series, a significant advancement over its predecessor, the GTE-Qwen series, in text embedding and reranking capabilities, built upon the Qwen3 foundation models. Leveraging the Qwen3 LLMs' robust capabilities in multilingual text understanding and generation, our innovative multi-stage training pipeline combines large-scale unsupervised pre-training with supervised fine-tuning on high-quality datasets. Effective model merging strategies further ensure the robustness and adaptability of the Qwen3 Embedding series. During the training process, the ","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":24},{"title":"Text Embeddings by Weakly-Supervised Contrastive Pre-training","work_id":"789cc674-467e-4f23-bb50-05c79fe8c4c2","shared_citers":17},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":15},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":12},{"title":"Embeddinggemma: Powerful and lightweight text representations","work_id":"e83995d5-78dc-4a7b-8d9a-1776d004351d","shared_citers":11},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":10},{"title":"Multilingual E5 Text Embeddings: A Technical Report","work_id":"b2004e20-1b24-4c37-9fad-dfdd9a7a8fee","shared_citers":9},{"title":"arXiv preprint arXiv:2312.02724 , year=","work_id":"0d9b3ad1-b405-412f-81ee-fd6f941d2367","shared_citers":8},{"title":"BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models","work_id":"c5f7f027-ac36-4b07-b824-0eca2f310641","shared_citers":8},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":8},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":8},{"title":"Retrieval-Augmented Generation for Large Language Models: A Survey","work_id":"b80d2790-6cd9-4c87-b3c4-de404f99a80e","shared_citers":8},{"title":"doi: 10.18653/v1/N19-1423","work_id":"3e3c8ac8-b858-4b22-af32-393d98c883e0","shared_citers":7},{"title":"Gemini embedding: Generalizable embeddings from gemini","work_id":"911b6918-a128-453f-ae99-94388c38fcb1","shared_citers":7},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":7},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":7},{"title":"The probabilistic relevance framework: BM25 and beyond.Foundations and Trends in Information Retrieval, 3(4):333–389","work_id":"3dfaa21d-3751-420b-84f7-aeceda058b63","shared_citers":7},{"title":"Towards General Text Embeddings with Multi-stage Contrastive Learning","work_id":"861a61de-66fe-49d1-b1ab-11f8b082a4cc","shared_citers":7},{"title":"doi: 10.18653/v1/ 2021.naacl-main.112","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"Mmteb: Massive multilingual text embedding benchmark","work_id":"774aa5f1-35ad-4b36-b6cc-5f461cfab347","shared_citers":6},{"title":"MS MARCO: A Human Generated MAchine Reading COmprehension Dataset","work_id":"78d498ce-11db-4f88-8eb0-40e0f86af615","shared_citers":6},{"title":"MTEB: Massive text embedding benchmark","work_id":"5c5a0bda-f984-4ea2-8938-cab73425e1b6","shared_citers":6}],"time_series":[{"n":111,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"3e0735af-daed-4e22-b48a-b7778e5e9a45","orcid":null,"display_name":"Baosong Yang","source":"manual","import_confidence":0.72},{"id":"6b268705-395a-45a3-8f13-05e372ec3623","orcid":null,"display_name":"Dingkun Long","source":"manual","import_confidence":0.72},{"id":"496bb56e-3efb-4d5a-936f-7267383c13cc","orcid":null,"display_name":"Huan Lin","source":"manual","import_confidence":0.72},{"id":"2ddce3fd-0f51-461c-8c9a-2b3ba42d0bd1","orcid":null,"display_name":"Mingxin Li","source":"manual","import_confidence":0.72},{"id":"6314c38d-420c-49be-9eff-c70b54f3eb8e","orcid":null,"display_name":"Xin Zhang","source":"manual","import_confidence":0.72},{"id":"8cfd545f-d741-4ca1-89e7-cc2394bfeca3","orcid":null,"display_name":"Yanzhao Zhang","source":"manual","import_confidence":0.72}]}}