{"work":{"id":"5c2060c6-427c-4321-be22-49ccae439d80","openalex_id":"https://openalex.org/W4415957649","doi":"10.48550/arxiv","arxiv_id":"2203.14987","raw_key":null,"title":"https://doi.org/10.48550/arXiv","authors":null,"authors_text":"Wagner GL, Silvestri S, Constantinou NC, et al (","year":2025,"venue":"cs.AI","abstract":"Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various languages. In this paper, we explore multilingual KG completion, which leverages limited seed alignment as a bridge, to embrace the collective knowledge from multiple languages. However, language alignment used in prior works is still not fully exploited: (1) alignment pairs are treated equally to maximally push parallel entities to be close, which ignores KG capacity inconsistency; (2) seed alignment is scarce and new alignment identification is usually in a noisily unsupervised manner. To tackle these issues, we propose a novel self-supervised adaptive graph alignment (SS-AGA) method. Specifically, SS-AGA fuses all KGs as a whole graph by regarding alignment as a new edge type. As such, information propagation and noise influence across KGs can be adaptively controlled via relation-aware attention weights. Meanwhile, SS-AGA features a new pair generator that dynamically captures potential alignment pairs in a self-supervised paradigm. Extensive experiments on both the public multilingual DBPedia KG and newly-created industrial multilingual E-commerce KG empirically demonstrate the effectiveness of SS-AG","external_url":"https://arxiv.org/abs/2203.14987","cited_by_count":24,"metadata_source":"doi_reference","metadata_fetched_at":"2026-07-09T21:46:34.506658+00:00","pith_arxiv_id":"2203.14987","created_at":"2026-05-08T12:34:59.002431+00:00","updated_at":"2026-07-09T21:46:34.506658+00:00","title_quality_ok":false,"display_title":"Dickerson","render_title":"Dickerson"},"hub":{"state":{"work_id":"5c2060c6-427c-4321-be22-49ccae439d80","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":681,"external_cited_by_count":24,"distinct_field_count":88,"first_pith_cited_at":"2022-07-18T21:01:17+00:00","last_pith_cited_at":"2026-07-08T09:21:22+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-19T14:09:40.213278+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":159},{"context_role":"method","n":19},{"context_role":"baseline","n":6},{"context_role":"dataset","n":2},{"context_role":"other","n":2}],"polarity_counts":[{"context_polarity":"background","n":141},{"context_polarity":"use_method","n":19},{"context_polarity":"unclear","n":11},{"context_polarity":"support","n":9},{"context_polarity":"baseline","n":6},{"context_polarity":"use_dataset","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"URLhttps://doi.org/10.48550/arXiv","claims":[{"claim_text":"Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various languages. In this paper, we explore multilingual KG completion, which leverages limited seed alignment as a bridge, to embrace the collective knowledge from multiple languages. However, language alignment used in prior works is still not fully exploited: (1) alignment pairs are treated equally to maximally push parallel entities to be close, which ignores KG capacity inconsistency; ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"5.1.1 Data Filtering. To reduce the presence of misinformation and biases, an intuitive approach involves the careful selection of high-quality pre-training data from reliable sources. In this way, we can ensure the factual correctness of data while also minimizing the introduction of social biases. As early as the advent of GPT-2, Radford et al. [252]underscored the significance of exclusively scraping web pages that had undergone rigorous curation and filtration by human experts. However, as p","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks URLhttps://doi.org/10.48550/arXiv because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (1 contexts).","role_counts":[{"n":1,"context_role":"background"}]},"error":null,"updated_at":"2026-05-14T18:36:37.646655+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"ec03d829-4569-4c6b-a60d-f9014d00c6a9","orcid":null,"display_name":"Kausthubh Chandramouli"},{"id":"420fa8b2-bf57-4e26-b09c-521adcab3e6f","orcid":null,"display_name":"Kelly Mae Allen"},{"id":"3f84749e-5912-4f6a-8fdf-990f09799afe","orcid":null,"display_name":"Christopher Mori"},{"id":"b7e30824-da90-4917-8240-947f6846f97b","orcid":null,"display_name":"Dror Baron"},{"id":"fa3b2fb8-a841-48cb-bc51-3c5c5d4f9b47","orcid":null,"display_name":"and Mário A"}]},"error":null,"updated_at":"2026-05-14T18:36:27.619280+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T18:29:32.187878+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":21},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":17},{"title":"doi: 10.18653/v1/ 2024.findings-acl.586","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":15},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":12},{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":10},{"title":"2023.Parallel Programming for Multicore and Cluster Systems(3 ed.)","work_id":"cf4c4e77-acaa-46b4-b066-ddf045165d05","shared_citers":9},{"title":"A ViLA: Asynchronous vision-language agent for streaming multimodal data interaction","work_id":"c63d2b15-197c-4d28-9923-1b35f7cc7b73","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":9},{"title":"NAACL-LONG.102","work_id":"557a574b-31bd-4f07-9bd0-47bc7c74bcd8","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"Attention Is All You Need","work_id":"baafb5a2-5272-43bc-932f-09fa9ffe5316","shared_citers":8},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":8},{"title":"OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":8},{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":7},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":7},{"title":"Topology- Agnostic Detection of Temporal Money Laundering Flows in Billion-Scale Transactions","work_id":"452f975d-5e1f-47e0-967f-db1ed9da0d80","shared_citers":7},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":6},{"title":"Hilbert’s sixth problem: derivation of fluid equations via Boltzmann’s kinetic theory","work_id":"677737da-b96c-44b1-ac13-a4fa03c705ef","shared_citers":6},{"title":"In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)","work_id":"5955f330-cfab-45b0-9ff5-4dd04bca728a","shared_citers":6},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":6},{"title":"Quantum Fingerprinting","work_id":"4deaf489-c81b-4322-bb0a-41188b0ad4db","shared_citers":6}],"time_series":[{"n":3,"year":2023},{"n":1,"year":2024},{"n":5,"year":2025},{"n":270,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T18:36:27.820643+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T18:30:08.848119+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"URLhttps://doi.org/10.48550/arXiv","claims":[{"claim_text":"Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various languages. In this paper, we explore multilingual KG completion, which leverages limited seed alignment as a bridge, to embrace the collective knowledge from multiple languages. However, language alignment used in prior works is still not fully exploited: (1) alignment pairs are treated equally to maximally push parallel entities to be close, which ignores KG capacity inconsistency; ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"5.1.1 Data Filtering. To reduce the presence of misinformation and biases, an intuitive approach involves the careful selection of high-quality pre-training data from reliable sources. In this way, we can ensure the factual correctness of data while also minimizing the introduction of social biases. As early as the advent of GPT-2, Radford et al. [252]underscored the significance of exclusively scraping web pages that had undergone rigorous curation and filtration by human experts. However, as p","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks URLhttps://doi.org/10.48550/arXiv because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (1 contexts).","role_counts":[{"n":1,"context_role":"background"}]},"error":null,"updated_at":"2026-05-14T18:30:02.359161+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"URLhttps://doi.org/10.48550/arXiv","claims":[{"claim_text":"Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various languages. In this paper, we explore multilingual KG completion, which leverages limited seed alignment as a bridge, to embrace the collective knowledge from multiple languages. However, language alignment used in prior works is still not fully exploited: (1) alignment pairs are treated equally to maximally push parallel entities to be close, which ignores KG capacity inconsistency; ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"5.1.1 Data Filtering. To reduce the presence of misinformation and biases, an intuitive approach involves the careful selection of high-quality pre-training data from reliable sources. In this way, we can ensure the factual correctness of data while also minimizing the introduction of social biases. As early as the advent of GPT-2, Radford et al. [252]underscored the significance of exclusively scraping web pages that had undergone rigorous curation and filtration by human experts. However, as p","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks URLhttps://doi.org/10.48550/arXiv because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (1 contexts).","role_counts":[{"n":1,"context_role":"background"}]},"error":null,"updated_at":"2026-05-14T18:29:46.305915+00:00"}},"summary":{"title":"URLhttps://doi.org/10.48550/arXiv","claims":[{"claim_text":"Predicting missing facts in a knowledge graph (KG) is crucial as modern KGs are far from complete. Due to labor-intensive human labeling, this phenomenon deteriorates when handling knowledge represented in various languages. In this paper, we explore multilingual KG completion, which leverages limited seed alignment as a bridge, to embrace the collective knowledge from multiple languages. However, language alignment used in prior works is still not fully exploited: (1) alignment pairs are treated equally to maximally push parallel entities to be close, which ignores KG capacity inconsistency; ","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"5.1.1 Data Filtering. To reduce the presence of misinformation and biases, an intuitive approach involves the careful selection of high-quality pre-training data from reliable sources. In this way, we can ensure the factual correctness of data while also minimizing the introduction of social biases. As early as the advent of GPT-2, Radford et al. [252]underscored the significance of exclusively scraping web pages that had undergone rigorous curation and filtration by human experts. However, as p","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks URLhttps://doi.org/10.48550/arXiv because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (1 contexts).","role_counts":[{"n":1,"context_role":"background"}]},"graph":{"co_cited":[{"title":"Pattern Recognition 127 (2022), 108611","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":21},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":17},{"title":"doi: 10.18653/v1/ 2024.findings-acl.586","work_id":"8d675bdd-79ca-48d6-9163-fc17ce0e8ece","shared_citers":15},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":12},{"title":"URL https://doi.org/10.1109/CVPR52733","work_id":"7efbc2dd-b0f2-4f71-bb1c-d2fcf110d805","shared_citers":10},{"title":"2023.Parallel Programming for Multicore and Cluster Systems(3 ed.)","work_id":"cf4c4e77-acaa-46b4-b066-ddf045165d05","shared_citers":9},{"title":"A ViLA: Asynchronous vision-language agent for streaming multimodal data interaction","work_id":"c63d2b15-197c-4d28-9923-1b35f7cc7b73","shared_citers":9},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"ImageBind One Embedding Space to Bind Them All","work_id":"b9701eca-d05e-4d2e-9045-6761df4ba175","shared_citers":9},{"title":"NAACL-LONG.102","work_id":"557a574b-31bd-4f07-9bd0-47bc7c74bcd8","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"Attention Is All You Need","work_id":"baafb5a2-5272-43bc-932f-09fa9ffe5316","shared_citers":8},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":8},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":8},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":8},{"title":"OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations","work_id":"d0e5199d-8907-47b1-905a-07ab8b623a4c","shared_citers":8},{"title":"Barron, Ben Mildenhall, Mehdi S","work_id":"0a23d1b7-bd56-43cc-8a80-7c43ce994e1e","shared_citers":7},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":7},{"title":"Topology- Agnostic Detection of Temporal Money Laundering Flows in Billion-Scale Transactions","work_id":"452f975d-5e1f-47e0-967f-db1ed9da0d80","shared_citers":7},{"title":"Ddp: Diffusion model for dense visual prediction","work_id":"b8a8bb9e-1d31-40e2-9cab-ae21e338dde6","shared_citers":6},{"title":"Hilbert’s sixth problem: derivation of fluid equations via Boltzmann’s kinetic theory","work_id":"677737da-b96c-44b1-ac13-a4fa03c705ef","shared_citers":6},{"title":"In2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)","work_id":"5955f330-cfab-45b0-9ff5-4dd04bca728a","shared_citers":6},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":6},{"title":"Quantum Fingerprinting","work_id":"4deaf489-c81b-4322-bb0a-41188b0ad4db","shared_citers":6}],"time_series":[{"n":3,"year":2023},{"n":1,"year":2024},{"n":5,"year":2025},{"n":270,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"fa3b2fb8-a841-48cb-bc51-3c5c5d4f9b47","orcid":null,"display_name":"and Mário A","source":"manual","import_confidence":0.72},{"id":"3f84749e-5912-4f6a-8fdf-990f09799afe","orcid":null,"display_name":"Christopher Mori","source":"manual","import_confidence":0.72},{"id":"b7e30824-da90-4917-8240-947f6846f97b","orcid":null,"display_name":"Dror Baron","source":"manual","import_confidence":0.72},{"id":"ec03d829-4569-4c6b-a60d-f9014d00c6a9","orcid":null,"display_name":"Kausthubh Chandramouli","source":"manual","import_confidence":0.72},{"id":"420fa8b2-bf57-4e26-b09c-521adcab3e6f","orcid":null,"display_name":"Kelly Mae Allen","source":"manual","import_confidence":0.72}]}}