{"work":{"id":"59edaa01-a696-45b3-9a08-5eae777a799e","openalex_id":"https://openalex.org/W1614298861","doi":"10.48550/arxiv.1301.3781","arxiv_id":"1301.3781","raw_key":null,"title":"Efficient Estimation of Word Representations in Vector Space","authors":null,"authors_text":"Tomas Mikolov, Kai Chen, Greg Corrado, Jeffrey Dean","year":2013,"venue":"cs.CL","abstract":"We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuring syntactic and semantic word similarities.","external_url":"https://arxiv.org/abs/1301.3781","cited_by_count":18147,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1301.3781","created_at":"2026-05-10T00:19:46.851708+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Efficient Estimation of Word Representations in Vector Space","render_title":"Efficient Estimation of Word Representations in Vector Space"},"hub":{"state":{"work_id":"59edaa01-a696-45b3-9a08-5eae777a799e","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":187,"external_cited_by_count":18147,"distinct_field_count":27,"first_pith_cited_at":"2013-12-21T03:36:08+00:00","last_pith_cited_at":"2026-07-09T13:49:34+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T03:19:39.013104+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":22},{"context_role":"method","n":5},{"context_role":"dataset","n":1}],"polarity_counts":[{"context_polarity":"background","n":21},{"context_polarity":"use_method","n":4},{"context_polarity":"support","n":2},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Efficient Estimation of Word Representations in Vector Space","claims":[{"claim_text":"We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuri","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"and missing details by leveraging contextual and external domain knowledge to automatically produce more structured, actionable, and high-quality bug reports. To achieve this, our work focuses exclusively on more complex and testable bug report examples on the Minecraft domain, where we utilize the Minecraft Wiki1 as our external knowledge source. We curated our dataset from the Mo- jira platform [20], which hosts over 450,000 bug reports related to Minecraft. A significant portion of the user-g","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[41] S. Ouyang, X. Zhu, Z. Xiao, M. Jiang, Y. Meng, and J. Han. Rast: Reasoning activation in llms via small-model transfer, 2025. URLhttps://arxiv.org/abs/2506.15710. [42] N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner. Steering llama 2 via contrastive activation addition, 2024. URLhttps://arxiv.org/abs/2312.06681. [43] K. Park, Y. J. Choe, and V. Veitch. The linear representation hypothesis and the geometry of large language models, 2024. URLhttps://arxiv.org/ab","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"latent semantic analysis, word embeddings [32, 33], and ultimately large language models [34-36]. At each stage, the underlying computational assumption has been the same: words have meanings that can be represented as fixed points in a real-valued vector space, and the task of a model is to learn where those points are and how they compose. The transformer architecture [37, 38] is the most successful embodiment of this program. Transformer-based large language models have largely succeeded in p","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"tion Agent of EditRefiner as the evaluation metric on EditFHF-15K, while following the respective evaluation protocols for the other benchmarks. For individual modules, following [4, 44], Percep- tual Agent is evaluated using CC, SIM, KLD, AUC-Judd, and NSS. Reasoning Agent is assessed via distortion classification accuracy and semantic alignment with ROUGE [28], METEOR [1], Word2Vec [36], and SimCSE [13]. Evaluation Agent is measured using SRCC, KRCC, and PLCC. 5.2 Comparison sesults 5.2.1 Over","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"algorithm next evaluatesinstitutional affiliation similarity. Two author names are matched if their affiliation strings achieve a normalized Levenshtein similarity ratio [60] above 0.6. They also match if one affiliation string contains the other after removing non-alphanumeric characters. The third criterion computes the cosinesimilarity between papersusing Word2Vec-based [29] document vectors constructed from titles, abstracts, and keywords. A pair is matched when this similarity exceeds a lan","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"capabilities across diverse actions. The core principle of zero-shot learning is to extract shared knowledge from prior information and transfer it from seen classes to unseen JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 3 classes. Zhang et al. [36] are the first to achieve ZS-TAD by encoding seen and unseen activities using Word2Vec [37], effectively capturing shared semantic information. They further enhance label embeddings by incorporating the CLIP text encoder [38], which leads","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Efficient Estimation of Word Representations in Vector Space because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (19 contexts).","role_counts":[{"n":19,"context_role":"background"},{"n":5,"context_role":"method"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-22T09:33:37.296460+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"a7595dc6-d1f0-4a7e-bbd2-7f3c0c87d66b","orcid":null,"display_name":"Tomas Mikolov"},{"id":"5f3d5790-83f8-4afe-9019-67d262e63c15","orcid":null,"display_name":"Kai Chen"},{"id":"d754a7ad-7b89-4bad-bb7f-48696fdd9669","orcid":null,"display_name":"Greg Corrado"},{"id":"adcd959a-c8d6-4a9a-9260-3b508cc21369","orcid":null,"display_name":"Jeffrey Dean"}]},"error":null,"updated_at":"2026-05-22T09:33:37.762826+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T10:18:43.142285+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":7},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations","work_id":"aedf7950-7c35-4e28-a32d-bec290f51669","shared_citers":5},{"title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","work_id":"756f9764-ecd6-4672-8043-b37c698c7ad2","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"Attention Is All You Need","work_id":"baafb5a2-5272-43bc-932f-09fa9ffe5316","shared_citers":4},{"title":"Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell","work_id":"efb3082c-4f47-4d65-b49b-c56ba744fbbf","shared_citers":4},{"title":"Graph Attention Networks","work_id":"7dd5bb04-b448-4f32-9719-5fc799641fd9","shared_citers":4},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":4},{"title":"Semi-Supervised Classification with Graph Convolutional Networks","work_id":"21fff118-807d-49cd-8229-f7087ba57b5d","shared_citers":4},{"title":"The Linear Representation Hypothesis and the Geometry of Large Language Models","work_id":"a7b44adc-f2c2-4420-a27d-8ade97dd3b75","shared_citers":4},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":3},{"title":"arXiv preprint arXiv:1802.05365 , year=","work_id":"dd973cba-647d-49d3-9d24-061b637bb0cd","shared_citers":3},{"title":"arXiv preprint arXiv:2501.16496 , year=","work_id":"f55f2189-55b1-4a1c-acfb-a5fa7bfa9e86","shared_citers":3},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"Available: https://doi.org/10.1145/3560815","work_id":"b3532db7-8a2b-425b-ba78-f6d6863b3a25","shared_citers":3},{"title":"BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions","work_id":"511eeb84-4b95-46d5-b14f-50da43f4f19f","shared_citers":3},{"title":"CodeBERT: A Pre-Trained Model for Programming and Natural Languages","work_id":"abd850e7-0a69-416d-b504-be33a55a8399","shared_citers":3},{"title":"Cross- lingual language model pretraining","work_id":"1b427a55-e10e-463d-b44e-7fe0fd535403","shared_citers":3},{"title":"Deep Learning Scaling is Predictable, Empirically","work_id":"3638ccb4-3a4f-460e-8b6f-867a65922801","shared_citers":3}],"time_series":[{"n":1,"year":2013},{"n":2,"year":2019},{"n":2,"year":2020},{"n":1,"year":2023},{"n":2,"year":2024},{"n":54,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T10:18:49.640677+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T10:18:51.966988+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Efficient Estimation of Word Representations in Vector Space","claims":[{"claim_text":"We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuri","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"and missing details by leveraging contextual and external domain knowledge to automatically produce more structured, actionable, and high-quality bug reports. To achieve this, our work focuses exclusively on more complex and testable bug report examples on the Minecraft domain, where we utilize the Minecraft Wiki1 as our external knowledge source. We curated our dataset from the Mo- jira platform [20], which hosts over 450,000 bug reports related to Minecraft. A significant portion of the user-g","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"[41] S. Ouyang, X. Zhu, Z. Xiao, M. Jiang, Y. Meng, and J. Han. Rast: Reasoning activation in llms via small-model transfer, 2025. URLhttps://arxiv.org/abs/2506.15710. [42] N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner. Steering llama 2 via contrastive activation addition, 2024. URLhttps://arxiv.org/abs/2312.06681. [43] K. Park, Y. J. Choe, and V. Veitch. The linear representation hypothesis and the geometry of large language models, 2024. URLhttps://arxiv.org/ab","claim_type":"background","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"latent semantic analysis, word embeddings [32, 33], and ultimately large language models [34-36]. At each stage, the underlying computational assumption has been the same: words have meanings that can be represented as fixed points in a real-valued vector space, and the task of a model is to learn where those points are and how they compose. The transformer architecture [37, 38] is the most successful embodiment of this program. Transformer-based large language models have largely succeeded in p","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"tion Agent of EditRefiner as the evaluation metric on EditFHF-15K, while following the respective evaluation protocols for the other benchmarks. For individual modules, following [4, 44], Percep- tual Agent is evaluated using CC, SIM, KLD, AUC-Judd, and NSS. Reasoning Agent is assessed via distortion classification accuracy and semantic alignment with ROUGE [28], METEOR [1], Word2Vec [36], and SimCSE [13]. Evaluation Agent is measured using SRCC, KRCC, and PLCC. 5.2 Comparison sesults 5.2.1 Over","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"algorithm next evaluatesinstitutional affiliation similarity. Two author names are matched if their affiliation strings achieve a normalized Levenshtein similarity ratio [60] above 0.6. They also match if one affiliation string contains the other after removing non-alphanumeric characters. The third criterion computes the cosinesimilarity between papersusing Word2Vec-based [29] document vectors constructed from titles, abstracts, and keywords. A pair is matched when this similarity exceeds a lan","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"capabilities across diverse actions. The core principle of zero-shot learning is to extract shared knowledge from prior information and transfer it from seen classes to unseen JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 3 classes. Zhang et al. [36] are the first to achieve ZS-TAD by encoding seen and unseen activities using Word2Vec [37], effectively capturing shared semantic information. They further enhance label embeddings by incorporating the CLIP text encoder [38], which leads","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Efficient Estimation of Word Representations in Vector Space because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (19 contexts).","role_counts":[{"n":19,"context_role":"background"},{"n":5,"context_role":"method"},{"n":1,"context_role":"dataset"}]},"error":null,"updated_at":"2026-05-22T09:33:37.767622+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Efficient Estimation of Word Representations in Vector Space","claims":[{"claim_text":"We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuri","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Efficient Estimation of Word Representations in Vector Space because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T10:18:45.082560+00:00"}},"summary":{"title":"Efficient Estimation of Word Representations in Vector Space","claims":[{"claim_text":"We propose two novel model architectures for computing continuous vector representations of words from very large data sets. The quality of these representations is measured in a word similarity task, and the results are compared to the previously best performing techniques based on different types of neural networks. We observe large improvements in accuracy at much lower computational cost, i.e. it takes less than a day to learn high quality word vectors from a 1.6 billion words data set. Furthermore, we show that these vectors provide state-of-the-art performance on our test set for measuri","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Efficient Estimation of Word Representations in Vector Space because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":7},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":6},{"title":"Representation Learning with Contrastive Predictive Coding","work_id":"7b08a1d4-d565-424e-9c86-6ef244b7b90a","shared_citers":6},{"title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations","work_id":"aedf7950-7c35-4e28-a32d-bec290f51669","shared_citers":5},{"title":"DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter","work_id":"756f9764-ecd6-4672-8043-b37c698c7ad2","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":5},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":5},{"title":"Attention Is All You Need","work_id":"baafb5a2-5272-43bc-932f-09fa9ffe5316","shared_citers":4},{"title":"Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell","work_id":"efb3082c-4f47-4d65-b49b-c56ba744fbbf","shared_citers":4},{"title":"Graph Attention Networks","work_id":"7dd5bb04-b448-4f32-9719-5fc799641fd9","shared_citers":4},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":4},{"title":"Semi-Supervised Classification with Graph Convolutional Networks","work_id":"21fff118-807d-49cd-8229-f7087ba57b5d","shared_citers":4},{"title":"The Linear Representation Hypothesis and the Geometry of Large Language Models","work_id":"a7b44adc-f2c2-4420-a27d-8ade97dd3b75","shared_citers":4},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":3},{"title":"arXiv preprint arXiv:1802.05365 , year=","work_id":"dd973cba-647d-49d3-9d24-061b637bb0cd","shared_citers":3},{"title":"arXiv preprint arXiv:2501.16496 , year=","work_id":"f55f2189-55b1-4a1c-acfb-a5fa7bfa9e86","shared_citers":3},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":3},{"title":"Available: https://doi.org/10.1145/3560815","work_id":"b3532db7-8a2b-425b-ba78-f6d6863b3a25","shared_citers":3},{"title":"BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions","work_id":"511eeb84-4b95-46d5-b14f-50da43f4f19f","shared_citers":3},{"title":"CodeBERT: A Pre-Trained Model for Programming and Natural Languages","work_id":"abd850e7-0a69-416d-b504-be33a55a8399","shared_citers":3},{"title":"Cross- lingual language model pretraining","work_id":"1b427a55-e10e-463d-b44e-7fe0fd535403","shared_citers":3},{"title":"Deep Learning Scaling is Predictable, Empirically","work_id":"3638ccb4-3a4f-460e-8b6f-867a65922801","shared_citers":3}],"time_series":[{"n":1,"year":2013},{"n":2,"year":2019},{"n":2,"year":2020},{"n":1,"year":2023},{"n":2,"year":2024},{"n":54,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"d754a7ad-7b89-4bad-bb7f-48696fdd9669","orcid":null,"display_name":"Greg Corrado","source":"manual","import_confidence":0.72},{"id":"adcd959a-c8d6-4a9a-9260-3b508cc21369","orcid":null,"display_name":"Jeffrey Dean","source":"manual","import_confidence":0.72},{"id":"5f3d5790-83f8-4afe-9019-67d262e63c15","orcid":null,"display_name":"Kai Chen","source":"manual","import_confidence":0.72},{"id":"a7595dc6-d1f0-4a7e-bbd2-7f3c0c87d66b","orcid":null,"display_name":"Tomas Mikolov","source":"manual","import_confidence":0.72}]}}