REVIEW 5 major objections 5 minor 67 references
MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read MALAMUTE, the first education-based multilingual cloze-style dataset, extracts 116,887 probes from 71 university textbooks and shows that current language models have large subject-specific knowledge gaps.
desk verdict MALAMUTE is a genuinely useful new resource, but its headline knowledge-gap numbers lean on a thin QC sample. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The dataset is itself the machinery. It rests on the hierarchy built into the source open textbook library: each book belongs to one of eight domains, each domain splits into subdomains (roughly one college course each), and each subdomain's glossary terms define concepts tied to specific textbook sections. For every concept term, the pipeline scrapes the paragraph containing the term, creates a paragraph-level prompt by masking all occurrences of the term and hiding other occurrences as [HIDDEN], and creates a sentence-level prompt by splitting the paragraph and keeping only sentences that contain the term. Regular-expression and part-of-speech filters remove prompts with verbs like 'describe' or 'solve', references to figures, digit labels, and short or deictic openings; the surviving 116,887 prompts form the probe set. Evaluation uses top-k accuracy over an entity list for masked language models and best sub-span accuracy across three prompt designs for causal language models.
What would settle it
Have three domain-expert annotators label a stratified random sample of, say, 2,000 prompts from each language and subdomain for whether the masked term is uniquely inferable, and compare the ambiguity rate across subdomains; if ambiguity varies strongly, the reported subdomain ranking of models could change after excluding ambiguous prompts.
Extended reading notes
Core claim
MALAMUTE is claimed to be the first education-based cloze-style knowledge probing dataset, and the first probing dataset to pair sentence-level and paragraph-level template-free prompts for the same concepts. Each prompt is made by taking a glossary term from the source textbooks' editorial structure and replacing it with [MASK]; paragraph-level prompts also hide other occurrences of the same term as [HIDDEN]. The paper reports that on this dataset current masked and causal language models show substantial knowledge gaps at subdomain level even when aggregate scores look acceptable, and that paragraph context consistently raises accuracy relative to sentence context across nearly all subdomains and model types, implying a higher estimated lower bound of knowledge than sentence-only probes give.
Load-bearing premise
The argument leans on the assumption that the automated extraction and filtering pipeline, checked on a 495-prompt English sample, produces prompts that unambiguously point to the intended concept in every subdomain and every language, even though that sample is about 0.4% of the dataset.
Editorial extensions
If this is right
- If MALAMUTE is representative of university curricula, then a model's aggregate benchmark score is not enough to certify classroom readiness; subdomain-level scores must be reported.
- Teachers and developers can use the domain-subdomain-concept hierarchy to identify specific courses a model can or cannot support before deployment.
- The consistent sentence-versus-paragraph gap implies knowledge estimates rise with context, so future benchmarks should evaluate at both granularities to approximate a model's true lower bound.
- The English-Spanish-Polish performance gap means deployment in multilingual classrooms will need language-specific validation, not just multilingual model labels.
- Because prompts are generated automatically from an expanding open textbook library, the dataset can be extended and updated as textbooks are revised.
Reading between the lines
- A testable extension: use MALAMUTE-style extraction on the same textbooks' end-of-chapter exercises or lecture slides to see whether cloze probes predict performance on application-style questions, not just definitional recall.
- The roughly 6-7% of prompts judged not to convey the intended concept could be concentrated in particular subdomains; weighting subdomain scores by per-concept ambiguity would make the granular comparisons more robust.
- Because the quality-control sample covered only English prompts, the Spanish and Polish subdomain scores rest on the assumption that the extraction pipeline behaves identically across languages; a multilingual annotation sample would test that.
- The dataset's entity-list-based top-k evaluation may underestimate masked models, since it restricts predictions to textbook terms that appear as labels anywhere in the collection; richer answer surfaces would change ranks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MALAMUTE, a multilingual cloze-style probing dataset constructed by extracting glossary-term definitions from 71 OpenStax textbooks and masking each term, producing sentence-level and paragraph-level prompts in English, Spanish, and Polish. The authors report 116,887 prompts organized into 8 domains, 55 subdomains, and 33,361 concepts, and they evaluate five masked and eight causal language models, reporting overall accuracy and subdomain-level breakdowns. The central claim is that MALAMUTE is the first education-based cloze-style benchmark and that current models, despite decent overall scores, show large subject-specific and language-specific knowledge gaps.
Significance. If the quality assumptions hold, MALAMUTE would be a valuable community resource: it is the first cloze-style benchmark tied to university-level educational content, it is template-free at the probe level, it provides both sentence-level and paragraph-level variants, and it ships with code and data. The evaluation is a clean measurement rather than a fitted model, so circularity concerns are limited; using the dataset's concept list as the MLM candidate set is a scoring choice rather than a derivation. The main value is as a granular diagnostic tool for educational deployment. The strength of the comparative conclusions, however, depends on prompt quality being uniform across subdomains and languages, and the current evidence for that uniformity is weak.
major comments (5)
- [Section 3.4 / Table 3] Table 3's #Concepts column sums to 50,008 (42,172 for English alone), not the 33,361 distinct concepts claimed in the Abstract and Section 3.4. Likewise, the #Prompts column sums to 117,340 (100,711 English), not 116,887 (100,258 English). The manuscript should define whether the table reports unique concepts/prompts or textbook-specific occurrences, and the totals should be reconciled. As written, the headline dataset statistics are internally inconsistent.
- [Appendix C / Section 5.4] The quality-control study samples 9 prompts per subdomain and is English-only. With n=9, a single ambiguous prompt changes the estimated specificity rate for that subdomain by 11 percentage points, so the aggregate 92.9%/93.9% specificity figures do not establish that subdomain-level scores are comparable. If the roughly 6-7% ambiguous prompts cluster in subdomains such as Calculus 3, Marketing, or Business Law, the reported knowledge gaps in Section 5.4 could be artifacts of prompt ambiguity rather than model knowledge. The authors should provide per-subdomain QC results with confidence intervals, expand the QC sample, or explicitly qualify subdomain comparisons as exploratory. The absence of any human QC for the Spanish and Polish portions also weakens the cross-lingual comparisons in Section 5.1.
- [Appendix B] The filtration heuristics in Appendix B remove prompts starting with 'this/that/we/also', prompts under five words, prompts referencing figures, prompts with action verbs, and parenthesized labels. These rules are content-sensitive: math definitions frequently start with 'we', and business or computer-science terms often carry parenthetical acronyms. The paper reports no per-subdomain removal counts and no comparison of filtered versus retained prompt difficulty. Without such analysis, subdomain-level performance differences in Section 5.4 may reflect differential filtration rather than model knowledge. The authors should report filtration rates per subdomain and test whether the main subdomain rankings are robust to alternative filter settings.
- [Abstract / Section 1] The Abstract and Introduction describe the probes as 'expert-written, peer-reviewed.' Section 3.2 shows that the prompts are extracted automatically from OpenStax textbook prose: the textbook authors selected the glossary terms, and the peer review applies to the textbooks, not to the masked prompts themselves. This wording overstates the human curation of the probes and should be corrected, for example to 'based on expert-written, peer-reviewed textbook definitions.'
- [Table 4 / Section 4.2] The reporting of CLM results is internally inconsistent. Table 4's note says causal LM results are shown as [prompt 1, prompt 2, prompt 3], but the GPT-4 family and Llama-3.1-405B cells contain single numbers, and Section 4.2 states that these models were evaluated with only the in-context prompt selected after a pilot on smaller models. Since Section 5.2 shows that prompt choice can change a subdomain score by 18 points (Spanish Business Statistics: 23.0% vs. 5.0%), the single-number presentation obscures the uncertainty in the headline model rankings. The authors should report which prompts contribute to each number and give per-prompt or variance information for all models.
minor comments (5)
- [Appendix A.1] Appendix A.1 is incomplete: it displays only 'pti' and does not specify the actual instruction text for Prompt 1, unlike Prompts 2 and 3.
- [References] The reference list contains placeholder links for Kalo (2022) and Nayak (2023); these should be replaced with full bibliographic entries.
- [Tables 2, 5-7] There are several typos, including 'Managarial Accounting' (Tables 5-7), 'Entrepeneurship' (Table 2), and 'IW A' (Table 2).
- [Table 3] The column headers 'Words (Pg.)' and 'Words (Sent.)' should be spelled out or defined in the caption for clarity.
- [Section 1] The phrase 'establishing a new higher lower-bound of knowledge' is confusing and should be reworded for precision.
Circularity Check
No significant circularity; the dataset is an external artifact and the model evaluations are measurements rather than derived predictions.
full rationale
MALAMUTE is a dataset-construction and benchmark-evaluation paper, not a derivation paper. The central claims are that the dataset exists, that it is built from OpenStax textbooks, and that language models score at certain levels on it. None of these claims reduces to its own inputs by construction. The prompts are generated from textbook terms through an automated pipeline plus filtering, with human quality control on a sample; the model accuracies are empirical measurements against those prompts. The MLM evaluation uses the dataset's own entity list as the candidate set for top-k ranking, following Shaier et al. (2024a), but this is a scoring protocol, not a fitted parameter or a prediction derived from the dataset; it does not force any particular accuracy value. The self-citations to Shaier et al. (2024a) for template-free probing and entity-ranking methodology are used as prior design choices, but they are not load-bearing in the sense of making the reported knowledge-gap findings true by definition; the findings depend on measured model outputs. The acknowledged limitation that only 495 English prompts received human QC, with Spanish and Polish portions unannotated, is a quality or validity concern about subdomain and cross-lingual comparisons, not a circularity concern. No equation, fitted parameter, or self-referential uniqueness argument is present. The paper is therefore self-contained as an empirical benchmark contribution, and the appropriate circularity score is minimal.
Assumptions & free parameters
assumptions (5)
- domain assumption OpenStax textbooks accurately represent university-level curriculum concepts across the eight covered domains.
- domain assumption Index terms in OpenStax textbooks correspond to discrete, testable concepts.
- ad hoc to paper The automated filtration rules (regex and POS tagging) preserve prompt quality and specificity.
- ad hoc to paper A 495-prompt sample (9 per subdomain) is representative of the full 116k-prompt dataset for quality control.
- domain assumption Model performance on MALAMUTE is a valid proxy for educational knowledge.
Cite this review
Pith. "Pith review of MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset." pith.science (2026). https://pith.science/paper/ZGN6N5NX
@misc{pith2026241210105,
author = {Pith},
title = {Pith review of: MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZGN6N5NX}},
note = {Machine review of arXiv:2412.10105}
}
read the original abstract
Language models (LMs) have excelled in various broad domains. However, to ensure their safe and effective integration into real-world educational settings, they must demonstrate proficiency in specific, granular areas of knowledge. Existing cloze-style benchmarks, commonly used to evaluate LMs' knowledge, have three major limitations. They: 1) do not cover the educational domain; 2) typically focus on low-complexity, generic knowledge or broad domains, which do not adequately assess the models' knowledge in specific subjects; and 3) often rely on templates that can bias model predictions. Here, we introduce MALAMUTE, a multilingual, template-free, and highly granular probing dataset comprising expert-written, peer-reviewed probes from 71 university-level textbooks across three languages (English, Spanish, and Polish). MALAMUTE is the first education-based cloze-style dataset. It covers eight domains, each with up to 14 subdomains, further broken down into concepts and concept-based prompts, totaling 33,361 university curriculum concepts and 116,887 prompts. MALAMUTE's fine granularity, educational focus, and inclusion of both sentence-level and paragraph-level prompts make it an ideal tool for evaluating LMs' course-related knowledge. Our evaluation of masked and causal LMs on MALAMUTE shows that despite overall proficiency, they have significant gaps in knowledge when examined closely on specific subjects, hindering their safe use in classrooms and underscoring the need for further development.
Figures
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[4]
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. https://arxiv.org/abs/1903.10676 Scibert: A pretrained language model for scientific text . Preprint, arXiv:1903.10676
arXiv 2019
-
[5]
Prabin Bhandari, Antonios Anastasopoulos, and Dieter Pfoser. 2023. https://doi.org/10.1145/3589132.3625625 Are large language models geospatially knowledgeable? In Proceedings of the 31st ACM International Conference on Advances in Geographic Information Systems, SIGSPATIAL '23, New York, NY, USA. Association for Computing Machinery
arXiv 2023
-
[6]
Zied Bouraoui, Jose Camacho-Collados, and Steven Schockaert. 2019. https://arxiv.org/abs/1911.12753 Inducing relational knowledge from bert . Preprint, arXiv:1911.12753
work page Pith review arXiv 2019
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[8]
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 a . https://doi.org/10.18653/v1/2021.acl-long.146 Knowledgeable or educated guess? revisiting language models as knowledge bases . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Confere...
-
[9]
Boxi Cao, Hongyu Lin, Xianpei Han, Le Sun, Lingyong Yan, Meng Liao, Tong Xue, and Jin Xu. 2021 b . https://arxiv.org/abs/2106.09231 Knowledgeable or educated guess? revisiting language models as knowledge bases . Preprint, arXiv:2106.09231
arXiv 2021
Show all 67 references
-
[10]
Jie Cao, Rachel Dickler, Marie Grace, Jeffrey B Bush, Alessandro Roncone, Leanne M Hirshfield, Marilyn A Walker, and Martha S Palmer. 2023. Designing an ai partner for jigsaw classrooms. In Workshop on Language-Based AI Agent Interaction with Children (AIAIC’2023)
2023
-
[11]
Ilias Chalkidis, Nicolas Garneau, Catalina Goanta, Daniel Katz, and Anders S gaard. 2023. https://doi.org/10.18653/v1/2023.acl-long.865 L e XF iles and L egal LAMA : Facilitating E nglish multinational legal language model development . In Proceedings of the 61st Annual Meetin...
2023 doi
-
[12]
Manuel Ciosici, Joe Cecil, Dong-Ho Lee, Alex Hedges, Marjorie Freedman, and Ralph Weischedel. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.493 Perhaps PTLM s should go to school -- a task to assess open book and closed book QA . In Proceedings of the 2021 Conference on Em...
2021 doi
-
[13]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...
2020 doi
-
[14]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[15]
Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W
Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022. https://doi.org/10.1162/tacl_a_00459 Time-aware language models as temporal knowledge bases . Transactions of the Association for Computational Linguistics,...
2022 doi
-
[16]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783
2024 arXiv
-
[17]
Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.438 EXAMS : A multi-subject high school examinations dataset for cross-lingual and multilingual question answering . In Proceed...
2020 doi
-
[18]
Qiyuan He, Yizhong Wang, and Wenya Wang. 2024. https://arxiv.org/abs/2402.14273 Can language models act as knowledge bases at scale? Preprint, arXiv:2402.14273
2024 arXiv
-
[19]
Robin Jia and Percy Liang. 2017. https://doi.org/10.18653/v1/D17-1215 Adversarial examples for evaluating reading comprehension systems . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 2021--2031, Copenhagen, Denmark. Associati...
2017 doi
-
[20]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[21]
Xu, Jun Araki, and Graham Neubig
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. https://arxiv.org/abs/1911.12543 How can we know what language models know? Preprint, arXiv:1911.12543
2020 arXiv
-
[22]
Gregor Jo s t, Viktor Taneski, and Sa s o Karakati c . 2024. The impact of large language models on programming education and student learning outcomes. Applied Sciences, 14(10):4115
2024
-
[23]
Khapra, and Pratyush Kumar
Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul N.C., Avik Bhattacharyya, Mitesh M. Khapra, and Pratyush Kumar. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.445 I ndic NLPS uite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual langua...
2020 doi
-
[24]
Jan-Christoph Kalo. 2022. https://www.akbc.ws/2022/assets/pdfs/15_kamel_knowledge_analysis_with_.pdf [link]
2022
-
[25]
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, pages 15696--15707. PMLR
2023
-
[26]
Nora Kassner, Philipp Dufter, and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.eacl-main.284 Multilingual LAMA : Investigating knowledge in multilingual pretrained language models . In Proceedings of the 16th Conference of the European Chapter of the Association...
2021 doi
-
[27]
Nora Kassner, Benno Krojer, and Hinrich Schütze. 2020. https://arxiv.org/abs/2006.10413 Are pretrained language models symbolic reasoners over knowledge? Preprint, arXiv:2006.10413
2020 arXiv
-
[28]
Nora Kassner and Hinrich Schütze. 2020. https://arxiv.org/abs/1911.03343 Negated and misprimed probes for pretrained language models: Birds can talk, but cannot fly . Preprint, arXiv:1911.03343
2020 arXiv
-
[29]
Majeed Kazemitabaar, Runlong Ye, Xiaoning Wang, Austin Zachary Henley, Paul Denny, Michelle Craig, and Tovi Grossman. 2024. Codeaid: Evaluating a classroom deployment of an llm-based programming assistant that balances student and educator needs. In Proceedings of the CHI Conf...
2024
-
[30]
Amr Keleg and Walid Magdy. 2023. https://doi.org/10.18653/v1/2023.findings-acl.389 DLAMA : A framework for curating culturally diverse facts for probing the knowledge of pretrained language models . In Findings of the Association for Computational Linguistics: ACL 2023, pages ...
2023 doi
-
[31]
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, et al. 2024. Arabicmmlu: Assessing massive multitask language understanding in arabic. arXiv preprint arXiv:2402.12840
2024 arXiv
-
[32]
Harsh Kumar, Ilya Musabirov, Mohi Reza, Jiakai Shi, Anastasia Kuzminykh, Joseph Jay Williams, and Michael Liut. 2023. Impact of guidance and interaction strategies for llm use on learner performance and perception. arXiv preprint arXiv:2310.13712
2023 arXiv
-
[33]
Daliang Li, Ankit Singh Rawat, Manzil Zaheer, Xin Wang, Michal Lukasik, Andreas Veit, Felix Yu, and Sanjiv Kumar. 2023 a . https://doi.org/10.18653/v1/2023.findings-acl.112 Large language models with controllable working memory . In Findings of the Association for Computationa...
2023 doi
-
[34]
Lei Li, Jingjing Xu, Qingxiu Dong, Ce Zheng, Xu Sun, Lingpeng Kong, and Qi Liu. 2023 b . https://doi.org/10.18653/v1/2023.emnlp-main.726 Can language models understand physical concepts? In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,...
2023 doi
-
[35]
Bill Yuchen Lin, Seyeon Lee, Rahul Khanna, and Xiang Ren. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.557 B irds have four legs?! N umer S ense: P robing N umerical C ommonsense K nowledge of P re- T rained L anguage M odels . In Proceedings of the 2020 Conference on Emp...
2020 doi
-
[36]
https://ai.meta.com/blog/meta-llama-3/ Introducing meta llama 3: The most capable openly available llm to date
Llama3. https://ai.meta.com/blog/meta-llama-3/ Introducing meta llama 3: The most capable openly available llm to date
-
[37]
Linhao Luo, Trang Vu, Dinh Phung, and Reza Haf. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.885 Systematic assessment of factual knowledge in large language models . In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 13272--13286, Singapo...
2023 doi
-
[38]
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Compu...
2023
-
[39]
Zaiqiao Meng, Fangyu Liu, Ehsan Shareghi, Yixuan Su, Charlotte Collins, and Nigel Collier. 2022. https://doi.org/10.18653/v1/2022.acl-long.329 Rewire-then-probe: A contrastive recipe for probing biomedical knowledge of pre-trained language models . In Proceedings of the 60th A...
2022 doi
-
[40]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 arXiv
-
[41]
Anmol Nayak. 2023. https://openreview.net/pdf?id=nZM9Wxu3vw [link]
2023
-
[42]
B Nye, Dillon Mee, and Mark G Core. 2023. Generative large language models for dialog-based tutoring: An early consideration of opportunities and concerns. In AIED Workshops
2023
-
[43]
OpenAI. 2023 a . https://openai.com/blog/chatgpt/ Chatgpt: Optimizing language models for dialogue
2023
-
[44]
OpenAI. 2023 b . https://arxiv.org/abs/2303.08774 Gpt-4 technical report . Preprint, arXiv:2303.08774
2023 arXiv
-
[45]
OpenStax . 2012. https://openstax.org/ Openstax . Accessed: 2024-06-09. Originally launched in 2012. Continuously updated
2012
-
[46]
Miller, and Sebastian Riedel
Fabio Petroni, Patrick Lewis, Aleksandra Piktus, Tim Rocktäschel, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel. 2020. https://arxiv.org/abs/2005.04611 How context affects language models' factual predictions . Preprint, arXiv:2005.04611
2020 arXiv
-
[47]
Fabio Petroni, Tim Rockt \"a schel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. https://doi.org/10.18653/v1/D19-1250 Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language P...
2019 doi
-
[48]
Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.437 How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages...
2020 doi
-
[49]
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020. https://arxiv.org/abs/1910.01108 Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter . Preprint, arXiv:1910.01108
2020 arXiv
-
[50]
Sagi Shaier, Kevin Bennett, Lawrence Hunter, and Katharina Kann. 2023 a . https://doi.org/10.18653/v1/2023.ijcnlp-main.36 Emerging challenges in personalized medicine: Assessing demographic effects on biomedical question answering systems . In Proceedings of the 13th Internati...
2023 doi
-
[51]
Sagi Shaier, Kevin Bennett, Lawrence Hunter, and Katharina von der Wense. 2024 a . https://aclanthology.org/2024.eacl-long.46 Comparing template-based and template-free language model probing . In Proceedings of the 18th Conference of the European Chapter of the Association fo...
2024
-
[52]
Sagi Shaier, Lawrence Hunter, and Katharina Kann. 2023 b . https://doi.org/10.18653/v1/2023.ijcnlp-short.13 Who are all the stochastic parrots imitating? they should tell us! In Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd C...
2023 doi
-
[53]
Sagi Shaier, Lawrence Hunter, and Katharina von der Wense. 2024 b . https://aclanthology.org/2024.eacl-long.47 Desiderata for the context use of question answering systems . In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Ling...
2024
-
[54]
Sagi Shaier, Lawrence Hunter, and Katharina Wense. 2024 c . https://doi.org/10.18653/v1/2024.findings-acl.491 It is not about what you say, it is about how you say it: A surprisingly simple approach for improving reading comprehension . In Findings of the Association for Compu...
2024 doi
-
[55]
Sagi Shaier, Ari Kobren, and Philip V. Ogren. 2024 d . https://doi.org/10.18653/v1/2024.emnlp-main.956 Adaptive question answering: Enhancing language model proficiency for addressing knowledge conflicts with source citations . In Proceedings of the 2024 Conference on Empirica...
2024 doi
-
[56]
Mujeen Sung, Jinhyuk Lee, Sean Yi, Minji Jeon, Sungdong Kim, and Jaewoo Kang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.388 Can language models be biomedical knowledge bases? In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pag...
2021 doi
-
[57]
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020. https://arxiv.org/abs/1912.13283 olmpics -- on what language model pre-training captures . Preprint, arXiv:1912.13283
2020 arXiv
-
[58]
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. https://arxiv.org/abs/2211.09085 Galactica: A large language model for science . Preprint, arXiv:2211.09085
2022 arXiv
-
[59]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023. https://arxiv.org/abs/2302.13971 Llama:...
2023 arXiv
-
[60]
Cunxiang Wang, Xiaoze Liu, Yuanhao Yue, Xiangru Tang, Tianhang Zhang, Cheng Jiayang, Yunzhi Yao, Wenyang Gao, Xuming Hu, Zehan Qi, Yidong Wang, Linyi Yang, Jindong Wang, Xing Xie, Zheng Zhang, and Yue Zhang. 2023. https://arxiv.org/abs/2310.07521 Survey on factuality in large ...
2023 arXiv
-
[61]
Ruben Weijers, Gabrielle Fidelis de Castilho, Jean-Fran c ois Godbout, Reihaneh Rabbany, and Kellin Pelrine. 2024. https://aclanthology.org/2024.personalize-1.10 Quantifying learning-style adaptation in effectiveness of LLM teaching . In Proceedings of the 1st Workshop on Pers...
2024
-
[62]
Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024 a . https://arxiv.org/abs/2305.13300 Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts . Preprint, arXiv:2305.13300
2024 arXiv
-
[63]
Wenjing Xie, Juxin Niu, Chun Jason Xue, and Nan Guan. 2024 b . Grade like a human: Rethinking automated assessment with large language models. arXiv preprint arXiv:2405.19694
2024 arXiv
-
[64]
Dongfang Xu, Peter Jansen, Jaycie Martin, Zhengnan Xie, Vikas Yadav, Harish Tayyar Madabushi, Oyvind Tafjord, and Peter Clark. 2020. https://aclanthology.org/2020.lrec-1.661 Multi-class hierarchical question classification for multiple choice science exams . In Proceedings of ...
2020
-
[65]
Michael Zhang and Eunsol Choi. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.586 S ituated QA : Incorporating extra-linguistic contexts into QA . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7371--7387, Online and Punta C...
2021 doi
-
[66]
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, and Minlie Huang. 2023. Large language models are not robust multiple choice selectors. In The Twelfth International Conference on Learning Representations
2023
-
[67]
Zexuan Zhong, Dan Friedman, and Danqi Chen. 2021. https://doi.org/10.18653/v1/2021.naacl-main.398 Factual probing is [ MASK ]: Learning vs. learning to recall . In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics...
2021 doi
-
[68]
Li Zhou, Taelin Karidi, Nicolas Garneau, Yong Cao, Wanlong Liu, Wenyu Chen, and Daniel Hershcovich. 2024. https://arxiv.org/abs/2404.06833 Does mapo tofu contain coffee? probing llms for food-related cultural knowledge . Preprint, arXiv:2404.06833
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.