{"id":"46d70faa-e168-4595-90c0-62cc36ab72f7","arxiv_id":"2506.03895","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Systematic comparison shows Wikipedia2Vec embeddings and concept-aware entity linkers give the best entity retrieval re-ranking on DBpedia-Entity V2.","lead":"This paper compares three families of knowledge graph embeddings and five entity linking tools for entity retrieval, and shows that the choice matters: text-aware embeddings (Wikipedia2Vec) and linkers that catch both named entities and general concepts work best. Readers building entity search systems can use these results to choose the embedding and linker combination that maximizes coverage and concept recall.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The interpolation weight λ in Eq. (6) is never reported and appears to vary across configurations, so the comparative NDCG gains in Tables 4 and 5 are not fully reproducible; this is the most load-bearing threat to the central claim.","rationale":"The reader's weakest assumption is the reliability of the new Radboud annotations, and that is a legitimate concern for the concept-related conclusions (RQ1). However, I see a more load-bearing problem that sits underneath nearly every quantitative claim in the paper: the interpolation weight λ in Eq. (6) is never reported. Table 4 and Table 5 are the evidence for the abstract's central assertion, and all of their re-ranking rows are linear interpolations between the BM25F-CA baseline and the embedding-based F score. The value of λ controls how much of the observed gain is attributed to the graph-embedding component, and §5.5 explicitly mentions a lower λ for one variant, confirming the parameter varies. Without knowing λ or the selection protocol, one cannot tell whether the differences between linkers or embeddings are due to the methods themselves or due to per-configuration tuning. This also undermines the reported t-test significances, because significance testing normally assumes the compared configurations were not selected on the same test data. The annotation concern is partly independent: even if λ were fixed, the Radboud ground truth could bias RQ1. But the λ problem is more fundamental because it affects all three RQs and the headline NDCG numbers. I agree with the reader's CONDITIONAL verdict: the paper is a useful empirical study with a clear experimental framework, but these missing details must be documented before the conclusions can be used as a benchmark. The proposed check is cheap: fix λ and see whether the rankings survive; if they do, the central claim is much stronger.","tokens_in":21418,"tokens_out":5443,"duration_ms":52255,"concrete_test":"Download the public repository (github.com/informagi/GEEER), extract the λ values used for each row, and rerun the main comparisons with a single fixed λ (e.g., 0.5) for all rows of Tables 4 and 5, reporting NDCG@10 and NDCG@100. If the relative ordering of entity linkers in Table 4 or of graph embeddings in Table 5 changes, or the significant differences disappear, the central comparative claims are confounded by λ; if the ordering is stable, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that linker and embedding choice significantly changes retrieval effectiveness rests on Tables 4 and 5, whose scores are computed as scoretotal = (1−λ)·BM25F-CA + λ·F (Eq. 6). The paper never states λ or the procedure used to choose it, and §5.5 refers to 'the lower value of λ' for the no-graph variant, showing that different configurations use different λ values. If λ is tuned per configuration on the test queries, the NDCG differences and t-test significances (p<0.05) are not valid comparisons; if λ is fixed but unreported, the experiments cannot be reproduced. Since Eq. (6) is a convex combination of a strong lexical baseline and the embedding signal, the observed rankings (Combined > ELQ > TagMe > SMAPH > Nordlys > REL in Table 4; Wikipedia2Vec > ComplEx/RDF2Vec in Table 5) could be driven by the interpolation weight rather than by properties of the linker or embedding. The annotation-quality concern is real, but it mainly affects the concept-specific RQ1 interpretation; missing λ affects every headline number and every significance claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies entity retrieval by re-ranking BM25F-CA results with graph-embedding-based similarity scores. It compares three embedding families (Wikipedia2Vec, RDF2Vec, ComplEx) and five entity linkers, introduces new concept-focused annotations for the DBpedia-Entity V2 collection, and evaluates combinations using NDCG@10/@100 with t-test significance. The central claim is that both the choice of graph embedding and the choice of entity linker significantly affect retrieval effectiveness, with Wikipedia2Vec and concept-aware linkers performing best, and that maximizing entity coverage is important.","tokens_in":21594,"tokens_out":4984,"duration_ms":43062,"significance":"The paper contributes a systematic empirical comparison with public code and annotations, and the finding that linker and embedding choices matter is plausible and practically relevant. However, the validity of the main comparisons rests on the unreported interpolation weight λ in Eq. (6) and on a single-annotator ground truth with moderate agreement; these issues must be resolved before the headline claims can be accepted.","major_comments":[{"comment":"The interpolation weight λ in Eq. (6) is never reported, and §5.5 refers to 'the lower value of λ' for the no-graph variant, implying that λ was adjusted per configuration. Since all headline numbers are computed from Eq. (6) as a convex combination of BM25F-CA and the embedding score F, an unreported and potentially configuration-specific λ makes the NDCG differences and the paired t-test significances in Tables 4 and 5 unverifiable and, if λ was tuned on the test queries, invalid. Please report the exact λ value(s) used for every configuration, state how they were chosen, and provide a sensitivity analysis of the main rankings to λ.","section":"Section 3.4, Eq. (6), Section 5.5, Tables 4 and 5"},{"comment":"The new Radboud annotations, created by a single expert annotator with Cohen's kappa 0.54–0.59 on a 50-query subset, are the sole basis for the conclusion that concept annotations are important and that TagMe, SMAPH, and ELQ are the best linkers (RQ1, Section 5.2). The moderate agreement and the fact that only one annotator labeled the full set leave open the possibility that the observed advantage of Radboud over Webis annotations in Table 4 is an artifact of idiosyncratic annotation decisions. Please provide the annotation guidelines, analyze the disagreement cases, and test whether the retrieval conclusions change when using only the 50 queries with multiple annotations or when using a more conservative evaluation that downweights uncertain annotations.","section":"Section 4.2 and Table 4"},{"comment":"The coherence threshold τ is set to 'the highest value at which none of the box plots had a first quartile equal to 0' on the same 295 queries whose coherence scores are then compared in Figure 2. This data-dependent choice on the evaluation set makes the reported coherence differences (e.g., Wikipedia2Vec versus RDF2Vec and ComplEx) partly circular, and the paper should either pre-specify τ or demonstrate that the conclusions are stable across a range of τ values.","section":"Section 5.3, Eq. (8)-(9), Figure 2"}],"minor_comments":[{"comment":"In the description of the TransE scoring function, the vector expression is written as ||− →eh + − →r − − →eh||, which equals ||− →r ||; it should presumably be ||− →eh + − →r − − →et||.","section":"Section 3.3"},{"comment":"The column header 'T otal' contains a stray space.","section":"Table 1"},{"comment":"The query is referred to as 'spring shoe canada' in the text but 'spring shoes canada' in Table 6; please make these consistent.","section":"Section 5.5 and Table 6"},{"comment":"The third block of Table 8 is labeled 'Concepts' while the main text calls the same condition 'Radboud annotations'; this inconsistency should be fixed.","section":"Appendix A, Table 8"},{"comment":"The capitalization of RDF2Vec varies between 'RDF2vec' and 'RDF2Vec'; please standardize.","section":"Section 4.4"},{"comment":"The sentence 'we obtained them from the Nordlys package Hasibi et al. (2017a)' is missing a comma before the citation.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical comparison but the λ reporting and annotation quality are essential to address. The journal's scope is appropriate. The authors should be encouraged to release the exact runs and parameter settings."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a useful, systematic comparison of graph embeddings and entity linkers for entity retrieval, and the new Radboud annotations that include concepts are a genuine reusable resource. The main empirical claim is undercut by a reproducibility gap: the interpolation weight λ in Eq. (6) is never reported, and §5.5's mention of \"the lower value of λ\" for the no-graph variant suggests it was not a single fixed constant. If λ was tuned per configuration on the test queries, the NDCG differences and the significance tests are not valid comparisons; if it was fixed but simply unreported, the experiments cannot be reproduced. This is the first thing a referee should ask for.\n\nWhat the paper does well: the breadth of the comparison is new, it uses standard external test collections with NDCG and significance tests, and the coverage analysis (Table 2) makes a useful point: embedding coverage alone can dominate the choice of embedding family. The authors show real effort in recovering missing entities via redirects and page ID matching, and the results with and without pagelinks are a nice sanity check. The writing is clear, and the framing as an extension of Gerritse et al. (2020) is honest.\n\nWhere it's soft: (1) λ, as above. (2) RQ3's claim that graph structure helps retrieval rests on coherence scores and a handful of query examples, not an aggregate retrieval comparison. The coherence analysis is suggestive but not the same as showing NDCG gains for with-graph vs without-graph embeddings. (3) The new annotations were made mainly by one expert, with moderate inter-annotator agreement (kappa 0.54–0.59). That is not fatal, but it should temper the strong conclusion that concepts are the key factor; a larger or more reliable annotation effort would help. (4) The coherence threshold τ is selected via grid search on the same data, which is a minor concern.\n\nThe central claim that linker and embedding choice matter probably holds up, but as written the paper doesn't fully support its own headline because of the λ gap. This is fixable, not fatal. The paper deserves a serious referee; I would accept it with the expectation that the authors disclose λ, report its sensitivity, and ideally add an aggregate retrieval comparison for the with/without graph variant. For anyone working on entity retrieval components, this is worth reading once the details are clarified.","headline":"Useful systematic comparison, but an undisclosed interpolation weight λ undercuts the headline empirical claim.","tokens_in":22183,"tokens_out":3780,"would_cite":true,"duration_ms":35107,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68P20"],"pacs":[],"model":"deepseek-v4-flash","headline":"Entity retrieval effectiveness depends critically on the choice of graph embedding and entity linker; the best results come from Wikipedia2Vec embeddings with concept-aware linkers and near-complete entity coverage.","keywords":["entity retrieval","graph embeddings","knowledge graph embeddings","entity linking","Wikipedia2Vec","reranking","DBpedia-Entity","concept annotation"],"falsifier":"Create a second, independently produced concept-inclusive annotation set for DBpedia-Entity V2 with at least two annotators and high inter-annotator agreement; if concept-aware linkers (TagMe, SMAPH, ELQ) no longer show a consistent NDCG advantage over named-entity-only linkers on this new ground truth, the claim that concepts are essential is an artifact of the first annotations.","tokens_in":1711,"feed_emoji":"🔗","tokens_out":5422,"duration_ms":99138,"temperature":0.7,"pith_summary":"The paper asks a practical question: when using graph embeddings to re-rank entity retrieval results, which embedding method and which entity linker should you pick? Across three embedding families (skip-gram, random-walk, transition-based) and five linkers, the authors find that both choices significantly change effectiveness. The best system combines Wikipedia2Vec embeddings—which encode both Wikipedia's link graph and entity text—with linkers that annotate general concepts as well as named entities, and it pays to cover as many entities as possible in the embedding graph. On the DBpedia-Entity V2 collection this configuration raises NDCG@10 from 0.461 to 0.498 over the BM25F-CA baseline. A reader should care because most prior work fixed one embedding and one linker, leaving the real drivers of performance unexamined.","feed_headline":"Concept-aware linkers and text-graph embeddings win entity retrieval","feed_subtitle":"Systematic comparison shows Wikipedia2Vec reranking beats BM25F-CA, lifting NDCG@10 from 0.461 to 0.498.","key_machinery":"The core mechanism is a two-stage reranking pipeline: an initial lexical model (BM25F-CA) produces candidate entities, and the candidates are rescored by the embedding-based similarity $F(e,E_q)=\\sum_{e_q\\in E_q} s(e_q)\\,\\cos(\\vec{e},\\vec{e_q})$, where $E_q$ is the set of entities a linker detects in the query, $s(e_q)$ the linker's confidence, and $\\vec{e}$ an entity embedding. This score is interpolated with the baseline score through $\\lambda \\in [0,1]$, and for linkers that emit multiple query interpretations the maximum over interpretations is taken. The comparison is made meaningful by Wikipedia2Vec's joint objective $L = L_w + L_e + L_a$, which learns word vectors, entity vectors from link structure, and entity vectors from anchor-text context in one shared space; ablating the link term $L_e$ removes most of the retrieval gain, while coherence scores computed under the cluster hypothesis explain why the combined embeddings rank relevant entities closer.","core_discovery":"The central claim is that the choice of graph embedding and the choice of entity linker each materially change entity retrieval effectiveness, and that the best results come from a specific combination: an embedding method that jointly models link structure and textual descriptions (Wikipedia2Vec), an entity linker that recognizes both concepts and named entities, and an embedding graph that contains almost all entities of the target collection. The authors support this with a systematic comparison on DBpedia-Entity V2: re-ranking BM25F-CA with Wikipedia2Vec embeddings yields NDCG@10 0.484 with TagMe, and the union of all five linkers reaches 0.498, while structure-only or walk-based embeddings (ComplEx, RDF2Vec) lag unless large page-link files are added. They further show that a version of Wikipedia2Vec trained without the link graph produces worse clusters and lower retrieval effectiveness, supporting the conclusion that graph structure itself is a source of the gain.","pith_inferences":["The coverage result suggests a cheap, general recipe for retrieval pipelines: train or select embeddings on recent knowledge-graph dumps with minimal entity-frequency filtering, and measure coverage against the target collection before doing any tuning.","The link-graph ablation implies that cluster-coherence diagnostics could substitute for expensive retrieval evaluations when choosing among candidate embeddings.","The union-of-linkers result indicates that combining outputs from several independent linkers is a low-effort way to push retrieval quality further, especially when linkers have complementary strengths.","If concept annotations are what drive the gains here, transformer-based entity rankers that already encode entity descriptions may show similar sensitivity; testing concept-aware versus named-entity-only annotations on those models is a natural next step."],"forward_implications":["Entity retrieval systems should prefer graph embeddings that combine link structure with entity text, such as Wikipedia2Vec, over structure-only or walk-based embeddings.","Entity linkers that annotate both named entities and general concepts should be used, since concept annotations are what separate the best linkers (TagMe, SMAPH, ELQ) from the rest.","Maximizing entity coverage in the embedding graph is worth the extra training cost; the paper raises coverage from 75% to 97.6% and ties missing embeddings directly to retrieval losses.","Using the union of several linkers beats any single linker, suggesting that ensemble linking is a practical way to improve reranking without new training.","For transition- and walk-based embeddings such as ComplEx and RDF2Vec, including page-link triples is required to reach competitive coverage and effectiveness."],"supporting_citations":[{"why":"Supplies the reranking framework and Wikipedia2Vec setup that this paper systematically extends across embeddings and linkers.","marker":"(Gerritse et al., 2020)"},{"why":"Defines Wikipedia2Vec, the skip-gram-based joint word-entity embedding that proves most effective in the comparison.","marker":"(Yamada et al., 2016)"},{"why":"Introduces ComplEx, the transition-based embedding baseline that lags behind unless page links are added.","marker":"(Trouillon et al., 2016)"},{"why":"Introduces RDF2Vec, the random-walk-based embedding baseline whose coverage and effectiveness depend on the pagelinks file.","marker":"(Ristoski and Paulheim, 2016)"},{"why":"Provides DBpedia-Entity V2, the test collection used for all entity retrieval experiments and relevance assessments.","marker":"(Hasibi et al., 2017b)"},{"why":"Provides the Webis query annotations, the named-entity-only ground truth that the paper contrasts with its new concept-inclusive annotations.","marker":"(Kasturia et al., 2022)"},{"why":"Provides TagMe, one of the three entity linkers that annotate concepts and named entities and perform best.","marker":"(Ferragina and Scaiella, 2010)"},{"why":"Provides SMAPH, the concept-aware query linker whose high F-measure on concepts predicts strong reranking performance.","marker":"(Cornolti et al., 2018)"},{"why":"Provides ELQ, the end-to-end question linker that also annotates concepts and ranks among the best-performing linkers.","marker":"(Li et al., 2020)"}],"fun_headline_variants":["Wikipedia2Vec and concept-aware linkers boost entity retrieval","Graph-text embeddings plus concept linking win entity search","Joint graph-text embeddings improve entity retrieval reranking","Concept-aware entity linking with graph embeddings lifts NDCG","Best entity retrieval: Wikipedia2Vec embeddings with TagMe linking"],"cache_read_input_tokens":24320,"weakest_assumption_plain":"The conclusion that concept annotations are key rests on a new set of query annotations produced for this study by a single expert, with only moderate agreement (Cohen's kappa 0.54–0.59) on a 50-query subset.","fun_headline_variants_meta":{"raw":{"variants":["Wikipedia2Vec and concept-aware linkers boost entity retrieval","Graph-text embeddings plus concept linking win entity search","Joint graph-text embeddings improve entity retrieval reranking","Concept-aware entity linking with graph embeddings lifts NDCG","Best entity retrieval: Wikipedia2Vec embeddings with TagMe linking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1323,"prompt_tokens":889,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":505,"completion_tokens_details":{"reasoning_tokens":356}},"tokens_in":505,"tokens_out":434,"duration_ms":4432,"temperature":1.0,"reasoning_tokens":356,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:53:04.243656+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Create a second, independently produced concept-inclusive annotation set for DBpedia-Entity V2 with at least two annotators and high inter-annotator agreement; if concept-aware linkers (TagMe, SMAPH, ELQ) no longer show a consistent NDCG advantage over named-entity-only linkers on this new ground truth, the claim that concepts are essential is an artifact of the first annotations.","supporting_citations":[{"cited_title":"Graph-embedding empowered entity retrieval","cited_arxiv_id":null,"evidence_quote":"Supplies the reranking framework and Wikipedia2Vec setup that this paper systematically extends across embeddings and linkers."},{"cited_title":"Joint learning of the embedding of words and entities for named entity disambiguation","cited_arxiv_id":null,"evidence_quote":"Defines Wikipedia2Vec, the skip-gram-based joint word-entity embedding that proves most effective in the comparison."},{"cited_title":"Complex embeddings for simple link prediction","cited_arxiv_id":null,"evidence_quote":"Introduces ComplEx, the transition-based embedding baseline that lags behind unless page links are added."},{"cited_title":"RDF2Vec: RDF graph embeddings for data mining","cited_arxiv_id":null,"evidence_quote":"Introduces RDF2Vec, the random-walk-based embedding baseline whose coverage and effectiveness depend on the pagelinks file."},{"cited_title":"Query Interpretations from Entity-Linked Segmentations","cited_arxiv_id":null,"evidence_quote":"Provides the Webis query annotations, the named-entity-only ground truth that the paper contrasts with its new concept-inclusive annotations."},{"cited_title":"Tagme: on-the-fly annotation of short text fragments (by Wikipedia entities)","cited_arxiv_id":null,"evidence_quote":"Provides TagMe, one of the three entity linkers that annotate concepts and named entities and perform best."},{"cited_title":"u d, and Hinrich Sch\\","cited_arxiv_id":null,"evidence_quote":"Provides SMAPH, the concept-aware query linker whose high F-measure on concepts predicts strong reranking performance."}],"review_version":1}