REVIEW 2 major objections 5 minor 78 references
Enhancing Transformers for Generalizable First-Order Logical Entailment
T0 review · 2 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A transformer encoder with relative positional encoding can outperform specialized knowledge-graph query-answering methods at first-order logical entailment, and the paper's TEGA architecture, with logic-aware attention and free-variable…
desk verdict Useful benchmark and architecture study, but the 'substantially improves generalizability' claim is not supported for knowledge shift on two of three datasets; needs variance reporting before it can be trusted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is LogiRPE, a logic-aware relative positional encoding that augments self-attention with trainable bias vectors $\beta^K_{ij}$ and $\beta^V_{ij}$ indexed by the ordered pair of token types $(t_i, t_j)$ and the relative distance $|i-j|$; the token types are parenthesis, entity, relation, conjunction, disjunction, and negation. The second mechanism is free-variable pooling, which forms the final query embedding by pooling the last-layer hidden states of all free-variable tokens rather than reading only the first token. Together these two biases constitute TEGA, the Transformer Encoder with Guided Attention, and the paper's ablations attribute the reported gains to both components.
What would settle it
Run the same 55-type benchmark with the original released implementations of BetaE, ConE, CQD, LMPNN, SQE-LSTM, kgTransformer, and Query2Triple on FB15k-237; if any of them matches or beats the RPE transformer and TEGA on out-of-distribution query-type MRR, the paper's claim that generic transformers outperform methods designed for this task would be refuted.
Extended reading notes
Core claim
The paper's central discovery, stated on its own terms, is that a generic transformer encoder with relative positional encoding performs first-order logical entailment over parameterized knowledge at least as well as, and usually better than, methods engineered specifically for knowledge-graph query answering. The paper further finds that the two inductive biases previously proposed for this task, adjacency-matrix masking and directed-distance encoding, improve generalization only when paired with absolute positional encoding and fail to help under relative positional encoding, which the paper shows to be the superior setting. Filling that gap, the proposed TEGA architecture adds a logic-aware relative positional encoding and free-variable pooling, and the paper reports consistent MRR improvements over transformer baselines on FB15k, FB15k-237, and NELL995 in both in-distribution and out-of-distribution settings; for example, on FB15k-237 TEGA raises MRR on seen query types from 54.8 to 64.3 and on unseen query types from 37.7 to 42.6.
Load-bearing premise
The paper's comparative conclusions depend on its reimplementations of the specialized baselines and of the two prior inductive biases being faithful to the original methods, and on the benchmark's unseen query types in fact being absent from training; if either fails, the claims that transformers outperform task-specific models and that earlier biases are ineffective under relative positional encoding could be artifacts.
Editorial extensions
If this is right
- Relative positional encoding should replace absolute positional encoding as the default choice for transformer-based complex query answering, because it is more robust to query permutation and much stronger on unseen query types.
- Existing graph-structure inductive biases such as adjacency-matrix masking and directed-distance encoding should not be transplanted onto RPE transformers, since they only help under absolute positional encoding and masking can hurt performance.
- The EFO parallel syntax is preferable to Lisp-like nested syntax for RPE transformers, because the relative distances between logically related tokens are more consistent in the parallel structure.
- TEGA's two components, LogiRPE and free-variable pooling, each contribute independently, and combining them gives the best in-distribution and out-of-distribution MRR across all three knowledge graphs.
- Pre-trained KG embeddings such as ComplEx and DistMult can improve transformer KGQA performance, while TransE embeddings do not; jointly learned embeddings remain a strong baseline.
Reading between the lines
- A natural testable extension is to replace the hand-assigned token-type labels in LogiRPE with learned or inferred labels; if the gains survive, the same logic-aware attention could transfer to natural-language reasoning where tokens are not pre-tagged.
- The paper's benchmark protocol of scoring every method on both concept shift (unseen knowledge) and covariate shift (unseen query types) could become a standard evaluation grid for future KGQA work, making generalization claims directly comparable across papers.
- The authors' own limitations, namely single-free-variable EFO queries, explicit labeling, and knowledge graphs no larger than FB15k, leave open whether TEGA's advantage persists when queries have multiple free variables, when logical operators are unseen, or at web scale; those are the most direct stress tests of the architecture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether transformer encoders perform first-order logical entailment over knowledge graphs when knowledge is parameterized in model weights. It contributes a new benchmark with 23 seen and 32 unseen query types across FB15k, FB15k-237, and NELL995, with OOD splits along a knowledge dimension (concept shift) and a query-type dimension (covariate shift). It benchmarks transformer variants with four positional encodings against five KGQA baselines, studies query syntax, token embeddings, and architecture choices, and proposes TEGA, which combines a logic-aware relative positional encoding (LogiRPE) and free-variable pooling. The paper reports that transformers are competitive with or outperform specialized KGQA methods and that TEGA substantially improves performance and generalizability over relative-positional-encoding transformers.
Significance. When the OOD splits are taken at face value, the benchmark is useful and the experimental design is clear: the two shifts are defined precisely in Section 3.1, the query-type coverage is broader than prior benchmarks, and the ablations isolate the contributions of LogiRPE and free-variable pooling. The release of code and data is a reproducible-practice strength, and the paper is honest about the dependence of LogiRPE on explicit token-type labeling. The main reason the result is not yet fully established is that the knowledge-generalization improvements behind the headline claim are small and are reported without any variance estimate, while the query-type-generalization claims rest on one-seed comparisons.
major comments (2)
- [Section 6.3, Table 7] The paper's central claim that TEGA 'substantially improved the performance and generalizability' is not supported for the knowledge dimension. In Table 7, the OOD(K) column, which is the concept-shift metric introduced in Section 3.1, shows TEGA gaining between 0.1 and 1.8 MRR points over Trans.+Relative PE across the three datasets and both query-type splits, e.g., FB15k-237 ID(Q) OOD(K) is 20.1 vs. 20.0, NELL995 OOD(Q) OOD(K) is 13.7 vs. 13.4, and FB15k ID(Q) OOD(K) is 34.1 vs. 32.3. The paper reports no error bars, no seeds, and no significance tests anywhere, so these differences cannot be distinguished from noise based on the information provided. Because knowledge generalization is one of the two distribution shifts that define 'generalizable first-order logical entailment' in this paper, this issue is load-bearing for the main claim. Please report multi-seed variance and significance tests, or restrict the 'substantially improves' claim to the ID(K) and OOD(Q) ID(K) columns where the gains are large.
- [Section 5.3, Table 6] The conclusion that the existing inductive biases Adjacency Matrices Masking and Directed Distance Encoding are ineffective under RPE rests entirely on the authors' reimplementations, but the paper reports no hyperparameters, tuning budget, or validation that these reimplementations reproduce the behavior of kgTransformer and Query2Triple. An under-tuned reimplementation would make the 'APE-specific biases are ineffective' finding, which motivates TEGA, an artifact of the implementation rather than a property of the methods. Please provide implementation details and, if possible, a sanity check on the original benchmarks, or weaken the claim to apply to these specific reimplementations.
minor comments (5)
- [Section 1 and Abstract] The phrase 'we prove ineffective in the RPE setting' overstates the evidence, which is an empirical comparison on a single dataset (Table 6); use 'show' or 'find' instead.
- [Table 5] The column header 'Random' is misleading: this baseline is not a random KGE but a randomly initialized embedding matrix trained jointly with the transformer. Rename it to 'Jointly learned' or clarify in the caption.
- [Appendix G, Eqs. (9)-(10)] The MRR formula is malformed; as written it is not the standard mean reciprocal rank. Please restate it as (1/|A|) times the sum of 1/rank(v) over v in A.
- [Section 6.1] The six-token taxonomy in LogiRPE is tied to the EFO syntax of this benchmark. The Limitations section acknowledges this, but the main text should state upfront that the method requires a token-type annotation for any new syntax.
- [Section 5.4] Table 6 shows that Directed Distance Encoding roughly preserves RPE performance (e.g., OOD(Q) OOD(K) is 15.9 vs. 15.8), so the summary that existing biases 'fail to improve' should be phrased as 'do not improve substantially' to avoid overstatement.
Circularity Check
No significant circularity: the OOD evaluations are genuinely held out and no prediction reduces to a fitted parameter or to a self-citation chain.
full rationale
The paper's central claims are empirical benchmark results, not derivations from fitted parameters. The OOD(K) and OOD(Q) splits are defined by held-out knowledge and unseen query types (Appendix G), so TEGA's reported gains are measured on data that the model did not train on. The mapping from concept/covariate shift to unobserved knowledge/unseen query types in Section 3 is an interpretive formalization, not a result that is assumed into the evaluation. The paper does cite prior work from the same research group, notably the EFO syntax (Yin et al., 2023b) and the SQE-LSTM baseline (Bai et al., 2023b), but these citations supply representation conventions and comparison methods, not the paper's conclusions. No uniqueness theorem, ansatz, or fitted quantity is imported to force the TEGA design, which is introduced directly with its own inductive biases (LogiRPE and free-variable pooling) and evaluated against RPE baselines. The limitations paragraph's admission that LogiRPE depends on explicit labeling and that only medium-scale KGs were tested is a scope restriction, not evidence of circularity. Overall, the derivation chain is self-contained with respect to the experimental claims, and concerns about baseline implementation fidelity or statistical significance are correctness risks, not circular reasoning.
Assumptions & free parameters
free parameters (2)
- LogiRPE token type taxonomy =
fixed set of six types: parenthesis, entity, relation, conjunction, disjunction, negation
- Free-variable pooling operation =
sum pooling
assumptions (4)
- domain assumption A knowledge graph triple (h,r,t) represents the atomic first-order sentence r(h,t), and KGQA is equivalent to evaluating EFO formulas with a free variable substituted by candidate entities.
- domain assumption Concept shift and covariate shift from OOD generalization map onto unobserved knowledge and unseen query types, respectively.
- domain assumption MRR computed on Aood = A(G_test,q) - A(G_valid,q) measures knowledge generalization.
- ad hoc to paper The authors' reimplementations of Adjacency Matrices Masking and Directed Distance Encoding are faithful to the original methods.
Cite this review
Pith. "Pith review of Enhancing Transformers for Generalizable First-Order Logical Entailment." pith.science (2026). https://pith.science/paper/X2PPZOOC
@misc{pith2026250100759,
author = {Pith},
title = {Pith review of: Enhancing Transformers for Generalizable First-Order Logical Entailment},
year = {2026},
howpublished = {\url{https://pith.science/paper/X2PPZOOC}},
note = {Machine review of arXiv:2501.00759}
}
read the original abstract
Transformers, as the fundamental deep learning architecture, have demonstrated great capability in reasoning. This paper studies the generalizable first-order logical reasoning ability of transformers with their parameterized knowledge and how to improve it. Transformers' capability of first-order reasoning is further captured by whether they can conduct first-order logical entailment, which is quantitatively measured by their performance in answering knowledge graph queries. We establish the connections between (1) two types of distribution shifts studied in out-of-distribution generalization and (2) unseen knowledge and query settings discussed in the task of knowledge graph query answering, which makes it possible to characterize the fine-grained generalizability. Results on our comprehensive dataset showed that transformers \textit{outperform} previous methods designed particularly for this task and provided detailed empirical evidence about the impact of the input query syntax, token embedding, and transformer architectures on their reasoning capability. Interestingly, our results revealed the mismatch of positional encoding and other design choices of transformer architectures in previous practices. Motivated by this, we propose TEGA, a logic-aware architecture that significantly improves the performance in generalizable first-order logical entailment.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Erik Arakelyan, Daniel Daza, Pasquale Minervini, and Michael Cochez. 2021. https://arxiv.org/abs/2011.03459 Complex query answering with neural link predictors . Preprint, arXiv:2011.03459
arXiv 2021
-
[4]
Jiaxin Bai, Wei Fan, Qi Hu, Qing Zong, Chunyang Li, Hong Ting Tsang, Hongyu Luo, Yauwai Yim, Haoyu Huang, Xiao Zhou, Feng Qin, Tianshi Zheng, Xi Peng, Xin Yao, Huiwen Yang, Leijie Wu, Yi Ji, Gong Zhang, Renhai Chen, and Yangqiu Song. 2025 a . https://arxiv.org/abs/2505.23628 Autoschemakg: Autonomous knowledge graph construction through dynamic schema indu...
arXiv 2025
-
[5]
Jiaxin Bai, Xin Liu, Weiqi Wang, Chen Luo, and Yangqiu Song. 2023 a . https://arxiv.org/abs/2305.19068 Complex query answering on eventuality knowledge graph with implicit logical constraints . Preprint, arXiv:2305.19068
work page Pith review arXiv 2023
-
[6]
Jiaxin Bai, Yicheng Wang, Tianshi Zheng, Yue Guo, Xin Liu, and Yangqiu Song. 2024. https://arxiv.org/abs/2312.15643 Advancing abductive reasoning in knowledge graphs through complex logical hypothesis generation . Preprint, arXiv:2312.15643
work page Pith review arXiv 2024
-
[7]
Jiaxin Bai, Zihao Wang, Yukun Zhou, Hang Yin, Weizhi Fei, Qi Hu, Zheye Deng, Jiayang Cheng, Tianshi Zheng, Hong Ting Tsang, Yisen Gao, Zhongwei Xie, Yufei Li, Lixin Fan, Binhang Yuan, Wei Wang, Lei Chen, Xiaofang Zhou, and Yangqiu Song. 2025 b . https://arxiv.org/abs/2501.14224 Top ten challenges towards agentic neural graph databases . Preprint, arXiv:2501.14224
arXiv 2025
-
[8]
Jiaxin Bai, Tianshi Zheng, and Yangqiu Song. 2023 b . https://openreview.net/forum?id=ERqGqZzSu5 Sequential query encoding for complex query answering on knowledge graphs . Transactions on Machine Learning Research
work page 2023
Show all 78 references
-
[9]
Yushi Bai, Xin Lv, Juanzi Li, and Lei Hou. 2023 c . https://arxiv.org/abs/2212.09567 Answering complex logical queries on knowledge graphs via query computation tree optimization . Preprint, arXiv:2212.09567
2023 arXiv
-
[10]
David G. T. Barrett, Felix Hill, Adam Santoro, Ari S. Morcos, and Timothy Lillicrap. 2018. https://arxiv.org/abs/1807.04225 Measuring abstract reasoning in neural networks . Preprint, arXiv:1807.04225
2018 arXiv
-
[11]
Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Scott Wen tau Yih, and Yejin Choi. 2020. https://arxiv.org/abs/1908.05739 Abductive commonsense reasoning . Preprint, arXiv:1908.05739
2020 arXiv
-
[12]
Bollacker, Colin Evans, Praveen K
Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor. 2008. https://api.semanticscholar.org/CorpusID:207167677 Freebase: a collaboratively created graph database for structuring human knowledge . In SIGMOD Conference
2008
-
[13]
Antoine Bordes, Nicolas Usunier, Alberto Garcia-Duran, Jason Weston, and Oksana Yakhnenko. 2013. https://proceedings.neurips.cc/paper_files/paper/2013/file/1cecc7a77928ca8133fa24680a88d2f9-Paper.pdf Translating embeddings for modeling multi-relational data . In Advances in Neu...
2013
-
[14]
Bowman, Gabor Angeli, Christopher Potts, and Christopher D
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015. https://arxiv.org/abs/1508.05326 A large annotated corpus for learning natural language inference . Preprint, arXiv:1508.05326
2015 arXiv
-
[15]
Hruschka, and Tom M
Andrew Carlson, Justin Betteridge, Bryan Kisiel, Burr Settles, Estevam R. Hruschka, and Tom M. Mitchell. 2010. Toward an architecture for never-ending language learning. In Proceedings of the Twenty-Fourth AAAI Conference on Artificial Intelligence, AAAI'10, page 1306–1313. AAAI Press
2010
-
[16]
Jianshu Chen. 2023. https://arxiv.org/abs/2302.09458 Learning language representations with logical inductive bias . Preprint, arXiv:2302.09458
2023 arXiv
-
[17]
Xuelu Chen, Ziniu Hu, and Yizhou Sun. 2022. https://arxiv.org/abs/2108.02390 Fuzzy logic based logical query answering on knowledge graphs . Preprint, arXiv:2108.02390
2022 arXiv
-
[18]
Peter Clark, Oyvind Tafjord, and Kyle Richardson. 2020. https://arxiv.org/abs/2002.05867 Transformers as soft reasoners over language . Preprint, arXiv:2002.05867
2020 arXiv
-
[19]
Hanjun Dai, Hui Li, Tian Tian, Xin Huang, Lin Wang, Jun Zhu, and Le Song. 2018. https://proceedings.mlr.press/v80/dai18b.html Adversarial attack on graph structured data . In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Mach...
2018
-
[20]
Mostafa Dehghani, Stephan Gouws, Oriol Vinyals, Jakob Uszkoreit, and Łukasz Kaiser. 2019. https://arxiv.org/abs/1807.03819 Universal transformers . Preprint, arXiv:1807.03819
2019 arXiv
-
[21]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://arxiv.org/abs/1810.04805 Bert: Pre-training of deep bidirectional transformers for language understanding . Preprint, arXiv:1810.04805
2019 arXiv
-
[22]
Wei Fan, Tianshi Zheng, Yiran Hu, Zheye Deng, Weiqi Wang, Baixuan Xu, Chunyang Li, Haoran Li, Weixing Shen, and Yangqiu Song. 2025. https://arxiv.org/abs/2505.14104 Legal rule induction: Towards generalizable principle discovery from analogous judicial precedents . Preprint, a...
2025 arXiv
-
[23]
Tianqing Fang, Zeming Chen, Yangqiu Song, and Antoine Bosselut. 2024. https://arxiv.org/abs/2403.07398 Complex reasoning over logical queries on commonsense knowledge graphs . Preprint, arXiv:2403.07398
2024 arXiv
-
[24]
Yisen Gao, Jiaxin Bai, Tianshi Zheng, Qingyun Sun, Ziwei Zhang, Jianxin Li, Yangqiu Song, and Xingcheng Fu. 2025. https://arxiv.org/abs/2505.20948 Controllable logical hypothesis generation for abductive reasoning in knowledge graphs . Preprint, arXiv:2505.20948
2025 arXiv
-
[25]
Hamilton, Payal Bajaj, Marinka Zitnik, Dan Jurafsky, and Jure Leskovec
William L. Hamilton, Payal Bajaj, Marinka Zitnik, Dan Jurafsky, and Jure Leskovec. 2019. https://arxiv.org/abs/1806.01445 Embedding logical queries on knowledge graphs . Preprint, arXiv:1806.01445
2019 arXiv
-
[26]
Fabbri, Wojciech Kryscinski, Xi Victoria Lin, Caiming Xiong, and Dragomir Radev
Simeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Luke Benson, Lucy Sun, Ekaterina Zubova, Yujie Qiao, Matthew Burtell, David Peng, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Shafiq Joty, A...
2022 arXiv
-
[27]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://arxiv.org/abs/2006.03654 Deberta: Decoding-enhanced bert with disentangled attention . Preprint, arXiv:2006.03654
2021 arXiv
-
[28]
Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021. https://arxiv.org/abs/2103.03874 Measuring mathematical problem solving with the math dataset . Preprint, arXiv:2103.03874
2021 arXiv
-
[29]
Sepp Hochreiter and J \"u rgen Schmidhuber. 1997. Long short-term memory. Neural computation, 9(8):1735--1780
1997
-
[30]
Yifan Hou, Jiaoda Li, Yu Fei, Alessandro Stolfo, Wangchunshu Zhou, Guangtao Zeng, Antoine Bosselut, and Mrinmaya Sachan. 2023. https://arxiv.org/abs/2310.14491 Towards a mechanistic interpretation of multi-step reasoning capabilities of language models . Preprint, arXiv:2310.14491
2023 arXiv
-
[31]
Bhushan Kotnis, Carolin Lawrence, and Mathias Niepert. 2021. https://doi.org/10.1609/aaai.v35i6.16630 Answering complex queries in knowledge graphs with bidirectional sequence encoders . Proceedings of the AAAI Conference on Artificial Intelligence, 35(6):4968--4977
2021 doi
-
[32]
Guillaume Lample and François Charton. 2019. https://arxiv.org/abs/1912.01412 Deep learning for symbolic mathematics . Preprint, arXiv:1912.01412
2019 arXiv
-
[33]
Chunyang Li, Weiqi Wang, Tianshi Zheng, and Yangqiu Song. 2025. https://arxiv.org/abs/2502.16169 Patterns over principles: The fragility of inductive reasoning in llms under noisy observations . Preprint, arXiv:2502.16169
2025 arXiv
-
[34]
Xiao Liu, Shiyu Zhao, Kai Su, Yukuo Cen, Jiezhong Qiu, Mengdi Zhang, Wei Wu, Yuxiao Dong, and Jie Tang. 2022. https://doi.org/10.1145/3534678.3539472 Mask and reason: Pre-training knowledge graph transformers for complex logical queries . In Proceedings of the 28th ACM SIGKDD ...
2022
-
[35]
David Marker. 2002. https://doi.org/10.1007/b98860 Model Theory: An Introduction . Springer-Verlag, New York
2002 doi
-
[36]
Jose G Moreno-Torres, Troy Raeder, Roc \' o Alaiz-Rodr \' guez, Nitesh V Chawla, and Francisco Herrera. 2012. A unifying view on dataset shift in classification. Pattern recognition, 45(1):521--530
2012
-
[37]
Shikhar Murty, Pratyusha Sharma, Jacob Andreas, and Christopher Manning. 2023. https://doi.org/10.18653/v1/2023.acl-short.38 Grokking of hierarchical structure in vanilla transformers . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics ...
2023 doi
-
[38]
Liangming Pan, Alon Albalak, Xinyi Wang, and William Yang Wang. 2023. https://arxiv.org/abs/2305.12295 Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning . Preprint, arXiv:2305.12295
2023 arXiv
-
[39]
Mihir Parmar, Nisarg Patel, Neeraj Varshney, Mutsumi Nakamura, Man Luo, Santosh Mashetty, Arindam Mitra, and Chitta Baral. 2024. https://arxiv.org/abs/2404.15522 Logicbench: Towards systematic evaluation of logical reasoning ability of large language models . Preprint, arXiv:2...
2024 arXiv
-
[40]
Stanislas Polu and Ilya Sutskever. 2020. https://arxiv.org/abs/2009.03393 Generative language modeling for automated theorem proving . Preprint, arXiv:2009.03393
2020 arXiv
-
[41]
Joaquin Quionero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil Lawrence. 2009. Dataset shift in machine learning
2009
-
[42]
Hongyu Ren, Mikhail Galkin, Michael Cochez, Zhaocheng Zhu, and Jure Leskovec. 2023. https://arxiv.org/abs/2303.14617 Neural graph reasoning: Complex logical query answering meets graph databases . Preprint, arXiv:2303.14617
2023 arXiv
-
[43]
Hongyu Ren, Weihua Hu, and Jure Leskovec. 2020. https://arxiv.org/abs/2002.05969 Query2box: Reasoning over knowledge graphs in vector space using box embeddings . Preprint, arXiv:2002.05969
2020 arXiv
-
[44]
Hongyu Ren and Jure Leskovec. 2020. https://arxiv.org/abs/2010.11465 Beta embeddings for multi-hop logical reasoning in knowledge graphs . Preprint, arXiv:2010.11465
2020 arXiv
-
[45]
Abulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi, Mehran Kazemi, Najoung Kim, and He He. 2023. Testing the general deductive reasoning capacity of large language models using ood examples. Advances in Neural Information Processing Systems, 36
2023
-
[46]
David Saxton, Edward Grefenstette, Felix Hill, and Pushmeet Kohli. 2019. https://arxiv.org/abs/1904.01557 Analysing mathematical reasoning abilities of neural models . Preprint, arXiv:1904.01557
2019 arXiv
-
[47]
Peter Shaw, Jakob Uszkoreit, and Ashish Vaswani. 2018. https://arxiv.org/abs/1803.02155 Self-attention with relative position representations . Preprint, arXiv:1803.02155
2018 arXiv
-
[48]
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. 2023. https://arxiv.org/abs/2104.09864 Roformer: Enhanced transformer with rotary position embedding . Preprint, arXiv:2104.09864
2023 arXiv
-
[49]
Arnold, Tania Bedrax-Weiss, Fernando Pereira, and William W
Haitian Sun, Andrew O. Arnold, Tania Bedrax-Weiss, Fernando Pereira, and William W. Cohen. 2021. https://arxiv.org/abs/2004.03658 Faithful embeddings for knowledge base queries . Preprint, arXiv:2004.03658
2021 arXiv
-
[50]
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna. 2015. https://arxiv.org/abs/1512.00567 Rethinking the inception architecture for computer vision . Preprint, arXiv:1512.00567
2015 arXiv
-
[51]
Shiro Takagi, Ryutaro Yamauchi, and Wataru Kumagai. 2023. https://arxiv.org/abs/2311.09706 Towards autonomous hypothesis verification via language models with minimal guidance . Preprint, arXiv:2311.09706
2023 arXiv
-
[52]
Jidong Tian, Yitian Li, Wenqing Chen, Liqiang Xiao, Hao He, and Yaohui Jin. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.303 Diagnosing the first-order logical reasoning ability through L ogic NLI . In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan...
2021 doi
-
[53]
Kristina Toutanova and Danqi Chen. 2015. https://doi.org/10.18653/v1/W15-4007 Observed versus latent features for knowledge base and text inference . In Proceedings of the 3rd Workshop on Continuous Vector Space Models and their Compositionality, pages 57--66, Beijing, China. ...
2015 doi
-
[54]
Théo Trouillon, Johannes Welbl, Sebastian Riedel, Éric Gaussier, and Guillaume Bouchard. 2016. https://arxiv.org/abs/1606.06357 Complex embeddings for simple link prediction . Preprint, arXiv:1606.06357
2016 arXiv
-
[55]
Hong Ting Tsang, Zihao Wang, and Yangqiu Song. 2025. https://arxiv.org/abs/2504.16537 Transformers for complex query answering over knowledge hypergraphs . Preprint, arXiv:2504.16537
2025 arXiv
-
[56]
V. Vapnik. 1991. https://proceedings.neurips.cc/paper_files/paper/1991/file/ff4d5fbbafdf976cfdc032e3bde78de5-Paper.pdf Principles of risk minimization for learning theory . In Advances in Neural Information Processing Systems, volume 4. Morgan-Kaufmann
1991
-
[57]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. https://arxiv.org/abs/1706.03762 Attention is all you need . Preprint, arXiv:1706.03762
2023 arXiv
-
[58]
Lean Wang, Lei Li, Damai Dai, Deli Chen, Hao Zhou, Fandong Meng, Jie Zhou, and Xu Sun. 2023 a . https://arxiv.org/abs/2305.14160 Label words are anchors: An information flow perspective for understanding in-context learning . Preprint, arXiv:2305.14160
2023 arXiv
-
[59]
Wong, and Simon See
Zihao Wang, Yangqiu Song, Ginny Y. Wong, and Simon See. 2023 b . https://arxiv.org/abs/2301.08859 Logical message passing networks with one-hop inference on atomic formulas . Preprint, arXiv:2301.08859
2023 arXiv
-
[60]
Zihao Wang, Hang Yin, and Yangqiu Song. 2021. https://arxiv.org/abs/2109.08925 Benchmarking the combinatorial generalizability of complex query answering on knowledge graphs . Preprint, arXiv:2109.08925
2021 arXiv
-
[61]
Adina Williams, Nikita Nangia, and Samuel R. Bowman. 2018. https://arxiv.org/abs/1704.05426 A broad-coverage challenge corpus for sentence understanding through inference . Preprint, arXiv:1704.05426
2018 arXiv
-
[62]
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2019. https://arxiv.org/abs/1810.00826 How powerful are graph neural networks? Preprint, arXiv:1810.00826
2019 arXiv
-
[63]
Yao Xu, Shizhu He, Cunguang Wang, Li Cai, Kang Liu, and Jun Zhao. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.761 Q uery2 T riple: Unified query encoding for answering diverse complex queries over knowledge graphs . In Findings of the Association for Computational Li...
2023 doi
-
[64]
Bishan Yang, Wen tau Yih, Xiaodong He, Jianfeng Gao, and Li Deng. 2015. https://arxiv.org/abs/1412.6575 Embedding entities and relations for learning and inference in knowledge bases . Preprint, arXiv:1412.6575
2015 arXiv
-
[65]
Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. https://arxiv.org/abs/1909.03193 Kg-bert: Bert for knowledge graph completion . Preprint, arXiv:1909.03193
2019 arXiv
-
[66]
Hang Yin, Zihao Wang, Weizhi Fei, and Yangqiu Song. 2023 a . https://arxiv.org/abs/2307.13701 EFO _ k -cqa: Towards knowledge graph complex query answering beyond set operation . Preprint, arXiv:2307.13701
2023 arXiv
-
[67]
Hang Yin, Zihao Wang, and Yangqiu Song. 2023 b . https://arxiv.org/abs/2304.07063 Rethinking complex queries on knowledge graphs with neural link predictors . Preprint, arXiv:2304.07063
2023 arXiv
-
[68]
Xiaoyu You, Beina Sheng, Daizong Ding, Mi Zhang, Xudong Pan, Min Yang, and Fuli Feng. 2023. https://doi.org/10.1145/3543507.3583203 Mass: Model-agnostic, semantic and stealthy data poisoning attack on knowledge graph embedding . In Proceedings of the ACM Web Conference 2023, W...
2023
-
[69]
Zhanqiu Zhang, Jie Wang, Jiajun Chen, Shuiwang Ji, and Feng Wu. 2021. https://arxiv.org/abs/2110.13715 Cone: Cone embeddings for multi-hop reasoning over knowledge graphs . Preprint, arXiv:2110.13715
2021 arXiv
-
[70]
Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Lorraine Li, and Alane Suhr
Wenting Zhao, Justin T Chiu, Jena D. Hwang, Faeze Brahman, Jack Hessel, Sanjiban Choudhury, Yejin Choi, Xiang Lorraine Li, and Alane Suhr. 2024. https://arxiv.org/abs/2311.08469 Uncommonsense reasoning: Abductive reasoning about uncommon situations . Preprint, arXiv:2311.08469
2024 arXiv
-
[71]
Tianshi Zheng, Jiaxin Bai, Yicheng Wang, Tianqing Fang, Yue Guo, Yauwai Yim, and Yangqiu Song. 2024. https://arxiv.org/abs/2407.20564 Clr-fact: Evaluating the complex logical reasoning capability of large language models over factual knowledge . Preprint, arXiv:2407.20564
2024 arXiv
-
[72]
Wong, and Simon See
Tianshi Zheng, Yixiang Chen, Chengxi Li, Chunyang Li, Qing Zong, Haochen Shi, Baixuan Xu, Yangqiu Song, Ginny Y. Wong, and Simon See. 2025 a . https://arxiv.org/abs/2504.05081 The curse of cot: On the limitations of chain-of-thought in in-context learning . Preprint, arXiv:2504.05081
2025
-
[73]
Wong, and Simon See
Tianshi Zheng, Jiayang Cheng, Chunyang Li, Haochen Shi, Zihao Wang, Jiaxin Bai, Yangqiu Song, Ginny Y. Wong, and Simon See. 2025 b . https://arxiv.org/abs/2502.11176 Logidynamics: Unraveling the dynamics of logical inference in large language model reasoning . Preprint, arXiv:...
2025
-
[74]
Tianshi Zheng, Zheye Deng, Hong Ting Tsang, Weiqi Wang, Jiaxin Bai, Zihao Wang, and Yangqiu Song. 2025 c . https://arxiv.org/abs/2505.13259 From automation to autonomy: A survey on large language models in scientific discovery . Preprint, arXiv:2505.13259
2025
-
[75]
Tianshi Zheng, Weihan Li, Jiaxin Bai, Weiqi Wang, and Yangqiu Song. 2025 d . https://arxiv.org/abs/2412.08985 Knowshiftqa: How robust are rag systems when textbook knowledge shifts in k-12 education? Preprint, arXiv:2412.08985
2025 arXiv
-
[76]
Zhaocheng Zhu, Mikhail Galkin, Zuobai Zhang, and Jian Tang. 2022. https://arxiv.org/abs/2205.10128 Neural-symbolic models for logical queries on knowledge graphs . Preprint, arXiv:2205.10128
2022 arXiv
-
[77]
Qing Zong, Zhaowei Wang, Tianshi Zheng, Xiyu Ren, and Yangqiu Song. 2025. https://arxiv.org/abs/2412.20251 Comparisonqa: Evaluating factuality robustness of llms through knowledge frequency control and uncertainty . Preprint, arXiv:2412.20251
2025 arXiv
-
[78]
Daniel Zügner, Amir Akbarnejad, and Stephan Günnemann. 2019. https://doi.org/10.24963/ijcai.2019/872 Adversarial attacks on neural networks for graph data . In Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19 , pages 6246--62...
2019 doi
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.