REVIEW 5 major objections 5 minor 1 cited by
Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Large language models, including small quantized open models, can identify broader, narrower, and same-as relations between research topics zero-shot with high accuracy; top proprietary score is 0.967 F1, top open 7B model 0.920 F1.
desk verdict A solid, honest benchmark paper whose headline F1 scores are likely real, but whose zero-shot interpretation needs a contamination test before the strong claims hold. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is IEEE-Rel-1K, a gold standard of 1,000 topic pairs sampled from the IEEE Thesaurus: 250 broader, 250 narrower, 250 validated same-as, and 250 unrelated pairs. The argument runs through four zero-shot prompting designs: standard one-way, standard two-way, chain-of-thought one-way, and chain-of-thought two-way. The two-way variants ask the model to classify the pair in both orders and then merge the answers with a fixed rule set (Section 4.3) that uses the inverse relationship between broader and narrower, the symmetry of same-as, and a tie-breaker based on the string length of the topic names. This combination of benchmark, prompting, and consistency rules is what lets the paper attribute the high F1 scores to reasoning rather than to model scale.
What would settle it
Check whether performance collapses on relationship pairs that could not have appeared in pretraining data, for example topics introduced after a model's training cutoff; if it does, the high F1 comes from memorization, not zero-shot reasoning.
Extended reading notes
Core claim
The paper claims that zero-shot LLM classification of broader, narrower, and same-as relations between engineering research topics is not only feasible but highly accurate, and that prompt design drives much of the performance. On the IEEE-Rel-1K benchmark, the proprietary Claude 3 Sonnet reaches 0.967 average F1 with simple one-way prompting, while the open, 8-bit-quantized Dolphin-Mistral-7B reaches 0.920 F1 with a two-way chain-of-thought strategy. The authors interpret these results as evidence that modern LLMs already encode enough scientific knowledge to support automated ontology construction, and that smaller, cheaper models can be tuned to near-proprietary performance by asking them to reason through the relationship twice, in both topic orders, and combining the answers with hand-written consistency rules.
Load-bearing premise
The whole result rests on the assumption that the models' high scores show genuine zero-shot understanding of topic relations, rather than memory of the IEEE Thesaurus from pretraining, and that the hand-written merging rules add no hidden test-set bias.
Editorial extensions
If this is right
- Automated or semi-automated pipelines can generate and update research-topic ontologies at a fraction of the manual cost, because the hard relation-labelling step can be delegated to zero-shot LLM inference.
- For a given model, switching from standard one-way prompting to two-way chain-of-thought prompting can improve average F1 by more than 0.2 points, so prompt engineering is a first-order lever for ontology quality.
- Quantized 7B open models can approach the accuracy of much larger proprietary systems, making on-premises or resource-limited ontology generation practical.
- The same four-way relation scheme generalizes to SKOS-style broader/narrower/same-as structures, so the approach could be reused for other knowledge organization systems.
- The best models still confuse same-as with hierarchical relations, so synonymy detection is the remaining bottleneck for fully automated ontology construction.
Reading between the lines
- A contamination test on topic pairs created after each model's training cutoff would separate memorized thesaurus structure from genuine zero-shot reasoning; the paper does not run one.
- The length-based tie-breaker in the two-way merging rules could be exploiting a regularity specific to IEEE-Thesaurus names; re-running on a benchmark where surface-form length is uncorrelated with the gold label would show whether the rule generalizes.
- If these results transfer, the bottleneck for automated ontologies shifts from relation detection to synonymy detection, the category where even the best models lose the most points.
- The same protocol could build analogous benchmarks from other controlled vocabularies, such as medicine or physics thesauri, to test whether the engineering finding is domain-general.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces IEEE-Rel-1K, a 1,000-pair gold standard sampled from the IEEE Thesaurus, and evaluates 17 LLMs (open, quantised, proprietary) on a four-way relation classification task (broader, narrower, same-as, other) under four zero-shot prompting strategies (standard one-way/two-way and CoT one-way/two-way). The central claim is that LLMs identify semantic topic relations with high accuracy in zero-shot settings—Claude 3 Sonnet reaches 0.967 F1, Dolphin-Mistral-7B 0.920 F1, and Mixtral-8x7B 0.847 F1—and that two-way CoT prompting yields the best results, enabling smaller quantised models to approach proprietary-model performance. The paper also provides an error analysis and releases code and data.
Significance. If the findings hold, the paper provides useful evidence that off-the-shelf LLMs can serve as a scalable first pass for ontology relation extraction in the engineering domain, and the released benchmark and codebase are valuable assets for future work. The strengths include the breadth of the model comparison (17 models, 4 strategies), the release of the gold standard and code, and the explicit error analysis. However, the headline claim of zero-shot semantic reasoning is currently undercut by the absence of a contamination analysis, by hand-crafted integration rules that may overfit the benchmark, and by the lack of statistical significance testing; the significance of the contribution is therefore conditional on these methodological gaps being closed.
major comments (5)
- [§3.2; §5.1] The central claim that the evaluated LLMs perform zero-shot semantic reasoning is not separated from memorization of the IEEE Thesaurus. IEEE-Rel-1K is extracted directly from IEEE Thesaurus v1.02 (July 2023) and several models (e.g., GPT-4-Turbo, Claude 3) have training data through late 2023, so the benchmark may overlap with pretraining material. The paper itself acknowledges this risk in §5.1 ('the superior performance of proprietary models suggests they may have encountered scientific publications during pre-training') and cites Kandpal et al. on memorization, but no experiment quantifies it. I ask for a concrete contamination analysis: for example, evaluate the same models on (i) topics drawn from a different thesaurus (e.g., MeSH or ACM CCS) and (ii) synthetic or semantically unrelated term pairs designed to defeat surface-form heuristics. Without such evidence, the abstract's claim of 'zero-shot reasoning' and the parity claim for quantised models are not supported.
- [§4.3] The two-way integration rules in §4.3 are hand-crafted and are applied directly to the test set without a validation split. Rules 3 and 4 use a length-based tie-breaker (len(tA) <= len(tB) selects broader; otherwise narrower) that is not justified and may implicitly encode a regularity of the benchmark rather than a general linguistic principle. The observed gains of two-way over one-way strategies (Tables 2, 4, 5) therefore depend on the specific rule design. The authors should report sensitivity to alternative integration rules (e.g., majority vote, always-broader, random tie-break, or a learned combination) and, ideally, select the rules on a held-out development set.
- [§5.5; Tables 1–4] All F1 scores are reported as point estimates without confidence intervals or significance tests. With 250 examples per class, a difference of 0.03 in class-level F1 is typically within sampling noise; with 1,000 examples, several of the model differences highlighted in §5.5 (e.g., sonnet 0.967 vs. gpt-4 0.948 in Table 1; dolphin-mistral 0.920 vs. dolphin-mistral-dpo 0.906 in Table 4) may not be statistically reliable. The paper should report bootstrap confidence intervals and pairwise significance tests (e.g., McNemar's test) for the central comparisons, and should state the number of independent runs and sampling temperature per model, none of which are currently given in §4.4.
- [§3.2] The construction of the 'other' class is weaker than the construction of the other classes. The 250 negative pairs were randomly generated and filtered only against thesaurus relations, but no human validation or inter-annotator agreement is reported for this class, even though the same-as class received manual expert selection. If some 'other' pairs are actually semantically related (or if some non-other relations are absent from the thesaurus and hence omitted from the positive classes), the measured precision and recall are biased in ways that affect all models and, in particular, the error analysis in §6. The authors should validate the negative pairs with at least two annotators and report agreement, or otherwise justify that random pairs from the IEEE Thesaurus are a safe source of negatives.
- [Abstract; §4.4] The abstract and Section 7 claim that smaller quantised models 'require significantly fewer computational resources' or are 'more scalable and cost-efficient,' but the paper reports no runtime, cost, energy, or throughput measurements. Section 4.4 lists the hardware and services (V100/L4 GPUs, Bedrock, OpenAI API) but does not quantify the resources consumed by each model. Without such measurements, the parity claim is not supported. I recommend adding at least wall-clock time, cost per 1,000 pairs, or another efficiency metric for the compared models.
minor comments (5)
- [§4.4; §5.4] There are several typos: 'Cluade 3 Haiku' should be 'Claude 3 Haiku', 'KoldbolAI' should be 'KoboldAI', and 'gtp-4' should be 'gpt-4'.
- [§1; Appendix B] 'IEEE Theasurus' is misspelled in the Introduction, and in Appendix B 'the model is sked to provide' should be 'the model is asked to provide'.
- [Appendices A and B] The prompt templates are described as 'engineered through various refinements,' but the paper does not report which refinements were explored or whether the final prompt was selected on a development set; this matters because prompt tuning on the test set would confound the comparison. At minimum, state that the prompt was fixed before evaluation and describe any prompt variants tested.
- [§5.5] The text states that 13 models improved and 4 declined overall but does not list which four declined; the reader must infer them from Table 5. Please state the four models explicitly.
- [§3.3.3] The exact API model identifiers or snapshot dates for the proprietary models (e.g., a specific GPT-4-Turbo version and Claude 3 Sonnet version) are not given; providing these would improve reproducibility and support future contamination assessments.
Circularity Check
No significant circularity: the benchmark is external and the central measurements are independent of the paper's construction choices.
full rationale
The paper's central claim is an empirical evaluation of seventeen LLMs on IEEE-Rel-1K, a gold standard derived from the external IEEE Thesaurus. The labels for broader, narrower, same-as, and other are taken from an outside resource, with manual validation for synonyms, rather than being generated by or fitted to the models under test. The four prompting strategies and the two-way combination rules are heuristics explicitly described in Section 4.3; they are not fitted parameters and the paper does not rename any fitted value as a prediction. The self-citations (e.g., CSO, Klink-2, the KOS survey) are contextual and do not carry the load of the main result. The paper itself raises the possibility that proprietary models encountered scientific content during pre-training (Section 5.1), but this is a data-contamination and validity concern, not a circular derivation by construction, and the review instructions require quoting a specific reduction before flagging circularity. No load-bearing step reduces to its inputs, so the correct finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Two-way integration rules including length tie-breaker
- Balanced class sizes in gold standard =
250 per class
assumptions (6)
- domain assumption IEEE Thesaurus is a valid ground truth for semantic relationships between engineering research topics.
- domain assumption The four relationship categories are mutually exclusive and exhaustive.
- domain assumption Randomly generated 'other' pairs, filtered only against IEEE Thesaurus relations, are genuinely unrelated.
- domain assumption Single-run zero-shot LLM outputs are representative.
- domain assumption LLM performance on IEEE-Rel-1K transfers to ontology generation in practice.
- domain assumption The manually curated same-as labels are accurate.
Cite this review
Pith. "Pith review of Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field." pith.science (2026). https://pith.science/paper/XWEB7MBA
@misc{pith2026241208258,
author = {Pith},
title = {Pith review of: Large Language Models for Scholarly Ontology Generation: An Extensive Analysis in the Engineering Field},
year = {2026},
howpublished = {\url{https://pith.science/paper/XWEB7MBA}},
note = {Machine review of arXiv:2412.08258}
}
read the original abstract
Ontologies of research topics are crucial for structuring scientific knowledge, enabling scientists to navigate vast amounts of research, and forming the backbone of intelligent systems such as search engines and recommendation systems. However, manual creation of these ontologies is expensive, slow, and often results in outdated and overly general representations. As a solution, researchers have been investigating ways to automate or semi-automate the process of generating these ontologies. This paper offers a comprehensive analysis of the ability of large language models (LLMs) to identify semantic relationships between different research topics, which is a critical step in the development of such ontologies. To this end, we developed a gold standard based on the IEEE Thesaurus to evaluate the task of identifying four types of relationships between pairs of topics: broader, narrower, same-as, and other. Our study evaluates the performance of seventeen LLMs, which differ in scale, accessibility (open vs. proprietary), and model type (full vs. quantised), while also assessing four zero-shot reasoning strategies. Several models have achieved outstanding results, including Mixtral-8x7B, Dolphin-Mistral-7B, and Claude 3 Sonnet, with F1-scores of 0.847, 0.920, and 0.967, respectively. Furthermore, our findings demonstrate that smaller, quantised models, when optimised through prompt engineering, can deliver performance comparable to much larger proprietary models, while requiring significantly fewer computational resources.
Figures
Forward citations
Cited by 1 Pith paper
-
Agentic Retrieval of Topics and Insights from Earnings Calls
An LLM agent extracts financial topics from earnings calls, builds a hierarchical topic ontology, and uses topic mention trends to flag rising and falling themes.
Reference graph
Works this paper leans on
-
[1]
A. Salatino, T. Aggarwal, A. Mannocci, F. Osborne, E. Motta, A survey of knowledge organization systems of research fields: Resources and challenges, Quantitative Science Studies (2025) 1–44 doi:10.1162/qss_a_ 00363. 23
doi:10.1162/qss_a_ 2025
-
[2]
E. Dunne, K. Hulek, Mathematics subject classification 2020, EMS Newsletter 2020–3 (115) (2020) 5–6. doi: 10.4171/news/115/2. URL http://dx.doi.org/10.4171/NEWS/115/2
-
[3]
C. E. Lipscomb, Medical subject headings (mesh), Bulletin of the Medical Library Association 88 (3) (2000) 265
2000
-
[4]
Rous, Major update to acm’s computing classification system, Communications of the ACM 55 (11) (2012) 12–12
B. Rous, Major update to acm’s computing classification system, Communications of the ACM 55 (11) (2012) 12–12
2012
-
[5]
Osborne, E
F. Osborne, E. Motta, P. Mulholland, Exploring scholarly data with rexplore, in: H. Alani, L. Kagal, A. Fokoue, P. Groth, C. Biemann, J. X. Parreira, L. Aroyo, N. Noy, C. Welty, K. Janowicz (Eds.), The Semantic Web – ISWC 2013, Springer Berlin Heidelberg, Berlin, Heidelberg, 2013, pp. 460–477
2013
-
[6]
J. Beel, B. Gipp, S. Langer, C. Breitinger, Paper recommender systems: a literature survey, International Journal on Digital Libraries 17 (2016) 305–338
2016
-
[7]
Gusenbauer, N
M. Gusenbauer, N. R. Haddaway, Which academic search systems are suitable for systematic reviews or meta- analyses? evaluating retrieval qualities of google scholar, pubmed, and 26 other resources, Research synthesis methods 11 (2) (2020) 181–217
2020
-
[8]
Meloni, S
A. Meloni, S. Angioni, A. Salatino, F. Osborne, D. R. Recupero, E. Motta, Integrating conversational agents and knowledge graphs within the scholarly domain, Ieee Access 11 (2023) 22468–22489
2023
Show all 115 references
-
[9]
Angioni, A
S. Angioni, A. Salatino, F. Osborne, D. R. Recupero, E. Motta, Aida: A knowledge graph about research dynamics in academia and industry, Quantitative Science Studies 2 (4) (2021) 1356–1398
2021
-
[10]
Osborne, E
F. Osborne, E. Motta, Klink-2: integrating multiple web sources to generate semantic topic networks, in: The Semantic Web-ISWC 2015: 14th International Semantic Web Conference, Bethlehem, PA, USA, October 11-15, 2015, Proceedings, Part I 14, Springer, 2015, pp. 408–424
2015
-
[11]
Wang, A.-L
D. Wang, A.-L. Barab ´asi, The Science of Science, Cambridge University Press, 2021
2021
-
[12]
Bornmann, R
L. Bornmann, R. Mutz, Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references, Journal of the Association for Information Science and Technology 66 (11) (2015) 2215–2222. doi:https://doi.org/10.1002/asi.23329
2015 doi
-
[13]
Sanderson, B
M. Sanderson, B. Croft, Deriving concept hierarchies from text, in: Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval, 1999, pp. 206–213
1999
-
[14]
Osborne, E
F. Osborne, E. Motta, Mining semantic relations between research areas, in: The Semantic Web–ISWC 2012: 11th International Semantic Web Conference, Boston, MA, USA, November 11-15, 2012, Proceedings, Part I 11, Springer, 2012, pp. 410–426
2012
-
[15]
K. Han, P. Yang, S. Mishra, J. Diesner, Wikicssh: extracting computer science subject headings from wikipedia, in: ADBIS, TPDL and EDA 2020 Common Workshops and Doctoral Consortium: International Workshops: DOING, MADEISD, SKG, BBIGAP, SIMPDA, AIMinScience 2020 and Doctoral Co...
2020
-
[16]
OpenAlex, Openalex: End-to-end process for topic classification, 2024
2024
-
[17]
A. A. Salatino, T. Thanapalasingam, A. Mannocci, A. Birukou, F. Osborne, E. Motta, The Computer Sci- ence Ontology: A Comprehensive Automatically-Generated Taxonomy of Research Areas, Data Intelligence 2 (3) (2020) 379–416. arXiv:https://direct.mit.edu/dint/article-pdf/2/3/379...
2020 doi
-
[18]
Kojima, S
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, Y . Iwasawa, Large language models are zero-shot reasoners (2023). arXiv:2205.11916. URL https://arxiv.org/abs/2205.11916
2023 arXiv
-
[19]
M. L. Zeng, Knowledge organization systems (kos), KO Knowledge Organization 35 (2-3) (2008) 160–182
2008
-
[20]
Hedden, Taxonomies and controlled vocabularies best practices for metadata, Journal of Digital Asset Man- agement 6 (2010) 279–284
H. Hedden, Taxonomies and controlled vocabularies best practices for metadata, Journal of Digital Asset Man- agement 6 (2010) 279–284. doi:10.1057/dam.2010.29
2010 doi
-
[21]
Zaharee, Building controlled vocabularies for metadata harmonization, Bulletin of the American Society for Information Science and Technology 39 (2) (2013) 39–42
M. Zaharee, Building controlled vocabularies for metadata harmonization, Bulletin of the American Society for Information Science and Technology 39 (2) (2013) 39–42. arXiv:https://asistdl.onlinelibrary. wiley.com/doi/pdf/10.1002/bult.2013.1720390211, doi:10.1002/bult.2013.1720...
2013
-
[22]
Guidelines for the construction, format, and management of monolingual controlled vocabularies, Standard, National Information Standards Organization, Baltimore, Maryland (Jul. 2010). doi:10.3789/ansi.niso. z39.19-2005R2010
2010 doi
-
[23]
thesauri and interoperability with othervocabularies
Information and documentation. thesauri and interoperability with othervocabularies. interoperability with other vocabularies, Standard, International Organization for Standardization (Mar. 2013)
2013
-
[24]
R. F. Rasch, The nature of taxonomy, Image: Journal of Nursing Scholarship 19 (3) (1987) 147–149
1987
-
[25]
T. R. Gruber, A translation approach to portable ontology specifications, Knowledge Acquisition 5 (2) (1993) 199–220. doi:https://doi.org/10.1006/knac.1993.1008. URL https://www.sciencedirect.com/science/article/pii/S1042814383710083
1993
-
[26]
M. R. Genesereth, N. J. Nilsson, Logical foundations of artificial intelligence, Morgan Kaufmann, 2012. doi: 10.1016/C2009-0-27551-9
2012 doi
-
[27]
A. A. Salatino, T. Thanapalasingam, A. Mannocci, F. Osborne, E. Motta, The computer science ontology: a large-scale taxonomy of research areas, in: The Semantic Web–ISWC 2018: 17th International Semantic Web Conference, Monterey, CA, USA, October 8–12, 2018, Proceedings, Part ...
2018
-
[28]
A. A. Salatino, F. Osborne, A. Birukou, E. Motta, Improving editorial workflow and metadata quality at springer nature, in: The Semantic Web–ISWC 2019: 18th International Semantic Web Conference, Auckland, New Zealand, October 26–30, 2019, Proceedings, Part II 18, Springer, 20...
2019
-
[29]
C. Peng, F. Xia, M. Naseriparsa, F. Osborne, Knowledge graphs: Opportunities and challenges, Artificial Intel- ligence Review 56 (11) (2023) 13071–13102
2023
-
[30]
F ¨arber, D
M. F ¨arber, D. Lamprecht, J. Krause, L. Aung, P. Haase, Semopenalex: The scientific landscape in 26 billion rdf triples, in: International Semantic Web Conference, Springer, 2023, pp. 94–112
2023
-
[31]
M. Y . Jaradeh, A. Oelen, K. E. Farfar, M. Prinz, J. D’Souza, G. Kismih ´ok, M. Stocker, S. Auer, Open research knowledge graph: next generation infrastructure for semantic scholarly knowledge, in: Proceedings of the 10th International Conference on Knowledge Capture, 2019, pp...
2019
-
[32]
Dess ´ı, F
D. Dess ´ı, F. Osborne, D. Reforgiato Recupero, D. Buscaldi, E. Motta, Cs-kg: A large-scale knowledge graph of research entities and claims in computer science, in: International Semantic Web Conference, Springer, 2022, pp. 678–696
2022
-
[33]
T. Kuhn, C. Chichester, M. Krauthammer, N. Queralt-Rosinach, R. Verborgh, G. Giannakopoulos, A.-C. N. Ngomo, R. Viglianti, M. Dumontier, Decentralized provenance-aware publishing with nanopublications, PeerJ Computer Science 2 (2016) e78
2016
-
[34]
Maedche, S
A. Maedche, S. Staab, Learning ontologies for the semantic web., in: SemWeb, 2001
2001
-
[35]
Cimiano, J
P. Cimiano, J. V ¨olker, Text2onto: A framework for ontology learning and data-driven change discovery, in: International conference on application of natural language to information systems, Springer, 2005, pp. 227– 238
2005
-
[36]
Navigli, P
R. Navigli, P. Velardi, A. Cucchiarelli, F. Neri, Quantitative and qualitative evaluation of the OntoLearn ontol- ogy learning system, in: COLING 2004: Proceedings of the 20th International Conference on Computational Linguistics, COLING, Geneva, Switzerland, 2004, pp. 1043–10...
2004
-
[37]
Velardi, S
P. Velardi, S. Faralli, R. Navigli, Ontolearn reloaded: A graph-based algorithm for taxonomy induction, Com- putational Linguistics 39 (3) (2013) 665–707
2013
-
[38]
Devlin, M.-W
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, arXiv preprint arXiv:1810.04805 (2018)
2018 arXiv
-
[39]
Grootendorst, Bertopic: Neural topic modeling with a class-based tf-idf procedure, arXiv preprint arXiv:2203.05794 (2022)
M. Grootendorst, Bertopic: Neural topic modeling with a class-based tf-idf procedure, arXiv preprint arXiv:2203.05794 (2022)
2022 arXiv
-
[40]
M. Le, S. Roller, L. Papaxanthos, D. Kiela, M. Nickel, Inferring concept hierarchies from text corpora via hyperbolic embeddings, arXiv preprint arXiv:1902.00913 (2019)
2019 arXiv
-
[41]
C. Chen, K. Lin, D. Klein, Constructing taxonomies from pretrained language models, arXiv preprint arXiv:2010.12813 (2020)
2020 arXiv
-
[42]
Canito, J
A. Canito, J. Corchado, G. Marreiros, A systematic review on time-constrained ontology evolution in predictive maintenance, Artificial Intelligence Review 55 (4) (2022) 3183–3211
2022
-
[43]
Osborne, E
F. Osborne, E. Motta, Pragmatic ontology evolution: reconciling user requirements and application performance, 25 in: The Semantic Web–ISWC 2018: 17th International Semantic Web Conference, Monterey, CA, USA, October 8–12, 2018, Proceedings, Part I 17, Springer, 2018, pp. 495–512
2018
-
[44]
Babaei Giglou, J
H. Babaei Giglou, J. D’Souza, S. Auer, Llms4ol: Large language models for ontology learning, in: International Semantic Web Conference, Springer, 2023, pp. 408–427
2023
-
[45]
M. J. Saeedizade, E. Blomqvist, Navigating ontology development with large language models, in: European Semantic Web Conference, Springer, 2024, pp. 143–161
2024
-
[46]
V . K. Kommineni, B. K¨onig-Ries, S. Samuel, From human experts to machines: An llm supported approach to ontology and knowledge graph construction, arXiv preprint arXiv:2403.08345 (2024)
2024 arXiv
-
[47]
Zhang, V
B. Zhang, V . A. Carriero, K. Schreiberhuber, S. Tsaneva, L. S. Gonz ´alez, J. Kim, J. de Berardinis, Ontochat: a framework for conversational ontology engineering using language models, in: European Semantic Web Con- ference, Springer, 2024, pp. 102–121
2024
-
[48]
Y . Sun, H. Xin, K. Sun, Y . E. Xu, X. Yang, X. L. Dong, N. Tang, L. Chen, Are large language models a good replacement of taxonomies?, arXiv preprint arXiv:2406.11131 (2024)
2024 arXiv
-
[49]
Z. Shen, H. Ma, K. Wang, A web-scale system for scientific knowledge exploration, arXiv preprint arXiv:1805.12216 (2018)
2018 arXiv
-
[50]
Cadeddu, A
A. Cadeddu, A. Chessa, V . De Leo, G. Fenu, E. Motta, F. Osborne, D. R. Recupero, A. Salatino, L. Secchi, A comparative analysis of knowledge injection strategies for large language models in the scholarly domain, Engineering Applications of Artificial Intelligence 133 (2024) 108166
2024
-
[51]
S. S. Khanal, P. Prasad, A. Alsadoon, A. Maag, A systematic review: machine learning based recommendation systems for e-learning, Education and Information Technologies 25 (4) (2020) 2635–2664
2020
-
[52]
Bolanos, A
F. Bolanos, A. Salatino, F. Osborne, E. Motta, Artificial intelligence for literature reviews: Opportunities and challenges, Artificial Intelligence Review 57 (10) (2024) 259
2024
-
[53]
Buscaldi, D
D. Buscaldi, D. Dess ´ı, E. Motta, M. Murgia, F. Osborne, D. R. Recupero, Citation prediction by leveraging transformers and natural language processing heuristics, Information Processing & Management 61 (1) (2024) 103583
2024
-
[54]
Brody, Scite, Journal of the Medical Library Association: JMLA 109 (4) (2021) 707
S. Brody, Scite, Journal of the Medical Library Association: JMLA 109 (4) (2021) 707
2021
-
[55]
Smith, Physics subject headings (physh), KO KNOWLEDGE ORGANIZATION 47 (3) (2020) 257–266
A. Smith, Physics subject headings (physh), KO KNOWLEDGE ORGANIZATION 47 (3) (2020) 257–266
2020
-
[56]
doi:10.3789/ansi.niso.z39.4-2021
Ansi /niso z39.4-2021, criteria for indexes (2021). doi:10.3789/ansi.niso.z39.4-2021. URL http://dx.doi.org/10.3789/ansi.niso.z39.4-2021
2021 doi
-
[57]
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, et al., Mistral 7b, arXiv preprint arXiv:2310.06825 (2023)
2023 arXiv
-
[58]
A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. d. l. Casas, E. B. Hanna, F. Bressand, et al., Mixtral of experts, arXiv preprint arXiv:2401.04088 (2024)
2024 arXiv
-
[59]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al., Llama 2: Open foundation and fine-tuned chat models, arXiv preprint arXiv:2307.09288 (2023)
2023 arXiv
-
[60]
Mukherjee, A
S. Mukherjee, A. Mitra, G. Jawahar, S. Agarwal, H. Palangi, A. Awadallah, Orca: Progressive learning from complex explanation traces of gpt-4 (2023). arXiv:2306.02707
2023 arXiv
-
[61]
Sharma, J
P. Sharma, J. T. Ash, D. Misra, The truth is in there: Improving reasoning in language models with layer-selective rank reduction, arXiv preprint arXiv:2312.13558 (2023)
2023 arXiv
-
[62]
Yadav, D
P. Yadav, D. Tam, L. Choshen, C. A. Ra ffel, M. Bansal, Ties-merging: Resolving interference when merging models, Advances in Neural Information Processing Systems 36 (2024)
2024
-
[63]
G. Wang, S. Cheng, X. Zhan, X. Li, S. Song, Y . Liu, Openchat: Advancing open-source language models with mixed-quality data, arXiv preprint arXiv:2309.11235 (2023)
2023 arXiv
-
[64]
de Bruin, J
T. de Bruin, J. Kober, K. Tuyls, R. Babu ˇska, Fine-tuning deep rl with gradient-free optimization, IFAC- PapersOnLine 53 (2) (2020) 8049–8056
2020
-
[65]
D. Kim, C. Park, S. Kim, W. Lee, W. Song, Y . Kim, H. Kim, Y . Kim, H. Lee, J. Kim, C. Ahn, S. Yang, S. Lee, H. Park, G. Gim, M. Cha, H. Lee, S. Kim, Solar 10.7b: Scaling large language models with simple yet effective depth up-scaling (2023). arXiv:2312.15166
2023 arXiv
-
[66]
W. Lian, B. Goodson, G. Wang, E. Pentland, A. Cook, C. V ong, ”Teknium”, Mistralorca: Mistral- 7b model instruct-tuned on filtered openorcav1 gpt-4 dataset, https://huggingface.co/Open-Orca/ 26 Mistral-7B-OpenOrca (2023)
2023
-
[67]
Longpre, L
S. Longpre, L. Hou, T. Vu, A. Webson, H. W. Chung, Y . Tay, D. Zhou, Q. V . Le, B. Zoph, J. Wei, A. Roberts, The flan collection: Designing data and methods for effective instruction tuning (2023). arXiv:2301.13688
2023 arXiv
-
[68]
L. Yuan, G. Cui, H. Wang, N. Ding, X. Wang, J. Deng, B. Shan, H. Chen, R. Xie, Y . Lin, et al., Advancing llm reasoning generalists with preference trees, arXiv preprint arXiv:2404.02078 (2024)
2024 arXiv
-
[69]
URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
AI@Meta, Llama 3 model card (2024). URL https://github.com/meta-llama/llama3/blob/main/MODEL_CARD.md
2024
-
[70]
Mitra, L
A. Mitra, L. D. Corro, S. Mahajan, A. Codas, C. Simoes, S. Agrawal, X. Chen, A. Razdaibiedina, E. Jones, K. Aggarwal, H. Palangi, G. Zheng, C. Rosset, H. Khanpour, A. Awadallah, Orca 2: Teaching small language models how to reason (2023). arXiv:2311.11045
2023 arXiv
-
[71]
J. Ye, X. Chen, N. Xu, C. Zu, Z. Shao, S. Liu, Y . Cui, Z. Zhou, C. Gong, Y . Shen, et al., A comprehensive capability analysis of gpt-3 and gpt-3.5 series models, arXiv preprint arXiv:2303.10420 (2023)
2023 arXiv
- [72]
-
[73]
Anthropic, The claude 3 model family: Opus, sonnet, haiku, online (2024)
2024
-
[74]
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, D. Zhou, Chain-of-thought prompting elicits reasoning in large language models (2023). arXiv:2201.11903. URL https://arxiv.org/abs/2201.11903
2023 arXiv
-
[75]
Kojima, S
T. Kojima, S. S. Gu, M. Reid, Y . Matsuo, Y . Iwasawa, Large language models are zero-shot reasoners, Advances in neural information processing systems 35 (2022) 22199–22213
2022
-
[76]
Y . Liu, Z. Guo, T. Liang, E. Shareghi, I. Vuli ´c, N. Collier, Aligning with logic: Measuring, evaluating and improving logical consistency in large language models, arXiv preprint arXiv:2410.02205 (2024)
2024 arXiv
-
[77]
Zhang, L
S. Zhang, L. Dong, X. Li, S. Zhang, X. Sun, S. Wang, J. Li, R. Hu, T. Zhang, F. Wu, et al., Instruction tuning for large language models: A survey, arXiv preprint arXiv:2308.10792 (2023)
2023
-
[78]
Bubeck, V
S. Bubeck, V . Chadrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y . T. Lee, Y . Li, S. Lundberg, et al., Sparks of artificial general intelligence: Early experiments with gpt-4 (2023)
2023
-
[79]
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong, et al., A survey of large language models, arXiv preprint arXiv:2303.18223 1 (2) (2023)
2023 arXiv
-
[80]
Kandpal, E
N. Kandpal, E. Wallace, C. Ra ffel, Deduplicating training data mitigates privacy risks in language models, in: International Conference on Machine Learning, PMLR, 2022, pp. 10697–10707
2022
-
[81]
Zadouri, A
T. Zadouri, A. ¨Ust¨un, A. Ahmadian, B. Ermis ¸, A. Locatelli, S. Hooker, Pushing mixture of experts to the limit: Extremely parameter efficient moe for instruction tuning, arXiv preprint arXiv:2309.05444 (2023)
2023 arXiv
-
[82]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, C. Finn, Direct preference optimization: Your language model is secretly a reward model, Advances in Neural Information Processing Systems 36 (2023) 53728–53741
2023
-
[83]
Z. Gero, C. Singh, H. Cheng, T. Naumann, M. Galley, J. Gao, H. Poon, Self-verification improves few-shot clinical information extraction, arXiv preprint arXiv:2306.00024 (2023)
2023 arXiv
-
[84]
X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, D. Zhou, Self-consistency improves chain of thought reasoning in language models, arXiv preprint arXiv:2203.11171 (2022)
2022 arXiv
-
[85]
R. Liu, J. Geng, A. J. Wu, I. Sucholutsky, T. Lombrozo, T. L. Gri ffiths, Mind your step (by step): Chain-of- thought can reduce performance on tasks where thinking makes humans worse, arXiv preprint arXiv:2410.21333 (2024)
2024 arXiv
-
[86]
Dettmers, M
T. Dettmers, M. Lewis, Y . Belkada, L. Zettlemoyer, Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale, Advances in neural information processing systems 35 (2022) 30318–30332
2022
-
[87]
Frantar, S
E. Frantar, S. Ashkboos, T. Hoefler, D.-A. Alistarh, Optq: Accurate post-training quantization for generative pre-trained transformers, in: 11th International Conference on Learning Representations, 2023
2023
-
[88]
Chang, X
Y . Chang, X. Wang, J. Wang, Y . Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y . Wang, et al., A survey on evaluation of large language models, ACM transactions on intelligent systems and technology 15 (3) (2024) 1–45
2024
-
[89]
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y . Mao, X. Ma, A. Efrat, P. Yu, L. Yu, et al., Lima: Less is more for alignment, Advances in Neural Information Processing Systems 36 (2023) 55006–55021
2023
-
[90]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al., Judging 27 llm-as-a-judge with mt-bench and chatbot arena, Advances in Neural Information Processing Systems 36 (2023) 46595–46623
2023
-
[91]
N. Miao, Y . W. Teh, T. Rainforth, Selfcheck: Using llms to zero-shot check their own step-by-step reasoning, arXiv preprint arXiv:2308.00436 (2023)
2023 arXiv
-
[92]
A. Pisu, L. Pompianu, A. Salatino, F. Osborne, D. Riboni, E. Motta, D. Reforgiato Recupero, et al., Leveraging language models for generating ontologies of research topics, in: CEUR WORKSHOP PROCEEDINGS, V ol. 3747, CEUR-WS, 2024, p. 11
2024
-
[93]
3759, CEUR, 2024
and others, Classifying scientific topic relationships with scibert, in: CEUR WORKSHOP PROCEEDINGS, V ol. 3759, CEUR, 2024
2024
-
[94]
Tsaneva, D
S. Tsaneva, D. Dess `ı, F. Osborne, M. Sabou, Knowledge graph validation by integrating llms and human-in-the- loop, Information Processing & Management (2025)
2025
-
[95]
Nenov, R
Y . Nenov, R. Piro, B. Motik, I. Horrocks, Z. Wu, J. Banerjee, Rdfox: A highly-scalable rdf store, in: The Semantic Web-ISWC 2015: 14th International Semantic Web Conference, Bethlehem, PA, USA, October 11- 15, 2015, Proceedings, Part II 14, Springer, 2015, pp. 3–20
2015
-
[96]
W. E. Marc ´ılio, D. M. Eler, From explanations to feature selection: assessing shap values as feature selection mechanism, in: 2020 33rd SIBGRAPI conference on Graphics, Patterns and Images (SIBGRAPI), Ieee, 2020, pp. 340–347
2020
-
[97]
J. Vig, A multiscale visualization of attention in the transformer model, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Association for Computa- tional Linguistics, Florence, Italy, 2019, pp. 37–42. doi:10.1...
2019 doi
-
[98]
Templeton, T
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, et al., Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet. transformer circuits thread (2024). 28 Appendix A. Prompt for Standard Prom...
2024
-
[101]
,→ ,→ ,→
'[TOPIC-A]' is-same-as-than '[TOPIC-B]' if '[TOPIC-A]' and '[TOPIC-B]' are synonymous terms denoting an identical concept (e.g., beautiful is-same-as-than attractive), including when one is the plural form of the other (e.g., cat is-same-as-than cats). ,→ ,→ ,→
-
[102]
,→ ,→ Given the previous definitions, determine which one of the following statements is correct:,→
'[TOPIC-A]' is-other-than '[TOPIC-B]' if '[TOPIC-A]' and '[TOPIC-B]' either have no direct relationship or share a different kind of relationship that does not fit into the other defined relationships. ,→ ,→ Given the previous definitions, determine which one of the following ...
-
[108]
Appendix B
'[TOPIC-A]' is-other-than '[TOPIC-B]' Answer by only stating the correct statement and its number. Appendix B. Prompt for Chain-of-Thought Prompting In the CoT prompting we have a two-phase interaction with the model. In the first interaction the model is sked to provide a def...
-
[109]
,→ ,→ ,→ ,→
'[TOPIC-A]' is-broader-than '[TOPIC-B]' if '[TOPIC-A]' is a super-category of '[TOPIC-B]', that is '[TOPIC-B]' is a type, a branch, or a specialised aspect of '[TOPIC-A]' or that '[TOPIC-B]' is a tool or a methodology mostly used in the context of '[TOPIC-A]' (e.g., car is-bro...
-
[110]
,→ ,→ ,→ ,→
'[TOPIC-A]' is-narrower-than '[TOPIC-B]' if '[TOPIC-A]' is a sub-category of '[TOPIC-B]', that is '[TOPIC-A]' is a type, a branch, or a specialised aspect of '[TOPIC-B]' or that '[TOPIC-A]' is a tool or a methodology mostly used in the context of '[TOPIC-B]' (e.g., wheel is-na...
-
[111]
,→ ,→ ,→
'[TOPIC-A]' is-same-as-than '[TOPIC-B]' if '[TOPIC-A]' and '[TOPIC-B]' are synonymous terms denoting a very similar concept (e.g., 'beautiful' is-same-as-than 'attractive'), including when one is the plural form of the other (e.g., cat is-same-as-than cats). ,→ ,→ ,→
-
[112]
,→ ,→ Think step by step by following these sequential instructions:
'[TOPIC-A]' is-other-than '[TOPIC-B]' if '[TOPIC-A]' and '[TOPIC-B]' either have no direct relationship or share a different kind of relationship that does not fit into the other defined relationships. ,→ ,→ Think step by step by following these sequential instructions:
-
[113]
Provide a precise definition for '[TOPIC-A]'
-
[114]
Provide a precise definition for '[TOPIC-B]'
-
[115]
Formulate a sentence that includes both '[TOPIC-A]' and '[TOPIC-B]'
-
[116]
Discuss '[TOPIC-A]' and '[TOPIC-B]' usage and relationship (is-narrower-than, is-broader-than, is-same-as-than, or is-other-than).,→ Second interaction [PREVIOUS-RESPONSE] Given the previous discussion, determine which one of the following statements is correct:,→
-
[117]
'[TOPIC-A]' is-broader-than '[TOPIC-B]'
-
[118]
'[TOPIC-B]' is-narrower-than '[TOPIC-A]'
-
[119]
'[TOPIC-A]' is-narrower-than '[TOPIC-B]'
-
[120]
'[TOPIC-B]' is-broader-than '[TOPIC-A]'
-
[121]
'[TOPIC-A]' is-same-as-than '[TOPIC-B]'
-
[122]
'[TOPIC-A]' is-other-than '[TOPIC-B]' Answer by only stating the number of the correct statement. 30
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.