REVIEW 4 major objections 6 minor 111 references
Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that small open-weight LLMs know ionic liquid facts but cannot reliably reason with them, as shown by a 5,920-example entailment benchmark where F1 collapses when the right answer is 'none of the above'.
desk verdict A genuinely useful expert-curated benchmark for LLM reasoning in the ionic-liquids domain, but the central 'reasoning gap' claim leans on a 'none-of-the-above' condition that may measure response bias rather than knowledge. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the entailment test bed itself: 5,920 expert-curated examples pairing a claim with candidate propositions, where the task is to select all entailing propositions or 'none of the above'. The dataset is constructed from 74 expert-written claims, 125 standardized universal propositions, incorrect variants at three difficulty levels (common-sense, mixed, expert-only), and paraphrased versions of both correct and incorrect options, arranged into 20 experiments across five groups that vary the number of adversarial options, the difficulty of wrong answers, and which options are paraphrased. The argumentative hinge is a consistency hypothesis: a knowledgeable agent should ignore added distracting options, pick 'none' when every option is false, and stay invariant under paraphrasing; the experiments are designed to expose reliance on linguistic cues when these expectations fail.
What would settle it
Re-run the Group 1 conditions with the 'none' option placed in different positions, or with explicit instructions that zero options may be correct; if median F1 returns to the 49-66 baseline, the collapse was a task-format artifact rather than evidence of absent reasoning.
Extended reading notes
Core claim
The paper's central claim is that smaller general-purpose LLMs hold basic factual knowledge about ionic liquids for carbon capture but cannot yet perform reliable domain-specific entailment reasoning. The evidence is a benchmark in which a claim is paired with candidate propositions and the model must choose every proposition that entails the claim, or 'none of the above' if none do. In the critical condition where every candidate is false, median F1 for Mistral and Gemma falls below 10, near zero in several settings, and to roughly 30 for Llama, against a baseline of 49 to 66 in the standard condition. The authors conclude that the models depend on linguistic and syntactic similarity rather than on chemistry knowledge, and that deploying them in carbon capture research without fine-tuning or augmentation would be unreliable.
Load-bearing premise
The central claim rests on assuming that a model failing to choose 'none of the above' when every proposition is false reflects missing domain reasoning, not a response bias such as always selecting at least one option; the authors themselves flag the position of the 'none' option as a possible confound.
Editorial extensions
If this is right
- If the claim is right, none of the three tested models should be deployed as an unsupervised reasoner in carbon capture research on ionic liquids.
- Fine-tuning on curated domain data, parameter-efficient methods such as LoRA, or retrieval-augmented generation become necessary steps before practical use.
- The released dataset gives the community a reusable benchmark that separates factual recall from applied reasoning in chemical and biological engineering.
- The sensitivity to paraphrasing and to the number of adversarial options implies that single-number accuracy on facts overstates what these models can do in niche domains.
- Aligning LLM development with carbon capture research is framed by the authors as a way to direct AI's own environmental cost toward climate solutions.
Reading between the lines
- The 'none-of-the-above' collapse may be a general property of small instruction-tuned models rather than a chemistry-specific gap; the same experiment design could be ported to other expert domains such as law or medicine.
- If response bias is the true driver, the Group 1 results measure decision threshold and calibration more than domain knowledge; a forced-choice variant with balanced option placement would separate the two.
- The three-tier difficulty design for incorrect propositions could be reused as a curriculum for fine-tuning, teaching models common-sense exclusions before expert-level ones.
- A panel of several domain experts would likely shift some difficulty labels, since a single non-expert evaluator agreed with the expert on factual correctness at only 67% F1.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces an expert-curated textual entailment dataset for evaluating large language models (LLMs) in the domain of ionic liquids (ILs) for carbon capture, and benchmarks three open-weight models (Llama-3.1-8B, Mistral-7B, Gemma-9B). The dataset comprises 5,920 examples built from 74 claims and 125 standardized propositions, with three difficulty levels of incorrect options, paraphrased variants, and varying option counts (5, 7, 10, 15), organized into 20 experiments. The authors report median F1 scores per experiment, finding that while all models achieve moderate baseline performance (median F1 roughly 49-66), performance collapses in Group 1 where all presented propositions are false and a 'none' option is available, with median F1 below 10 for Mistral and Gemma and around 30 for Llama. They interpret this drop as evidence that smaller LLMs possess IL-related factual knowledge but lack specialized domain reasoning, and discuss fine-tuning and augmentation strategies. The paper also analyzes the effects of the number of options, difficulty level of distractors, and paraphrasing on model performance.
Significance. The paper makes a useful empirical contribution by releasing a public, expert-curated benchmark (5,920 examples) for a niche scientific domain, with controlled difficulty and paraphrase perturbations. The evaluation protocol is transparent: deterministic generation at temperature 0, a structured prompt, and automated response formatting. The dataset is a genuinely novel resource that can support future work on domain-specific evaluation of LLMs in chemical and biological engineering. If the main finding holds, it would indicate that current small open-weight LLMs cannot reliably perform entailment reasoning over IL claims, which has practical implications for deployment in carbon capture research. The strength of the paper is its dataset and controlled experimental design; the main weakness is the construct validity of the Group 1 'none' condition as a measure of knowledge rather than response bias, a confound the authors themselves acknowledge. The paper's headline claim is plausible but not yet fully established.
major comments (4)
- [Section 4, 'Effect of only incorrect propositions as options'] The central inference from the Group 1 experiments is underdetermined by a response-bias confound. The paper interprets the dramatic F1 drop when only false propositions are offered as evidence that the models lack specialized reasoning, but this requires the assumption that selecting 'none' is a clean, unconfounded indicator of recognizing all options as false. The authors themselves note that the position of the 'none' option might be a confounding variable and defer the issue to future work. A model that is generally biased against selecting a 'none' rejection option, for reasons unrelated to IL knowledge, would produce the same collapse on any all-false item set. The paper should include an explicit control, for example all-false items drawn from general knowledge or a manipulated 'none' position/format, to rule out option-selection bias before attributing Group 1 results to missing domain reasoning.
- [Section 3.3, Phase 5] The paraphrase manipulation is load-bearing for Groups 3-5, but the paper does not report any verification that the Llama-3.1-8B-generated paraphrases preserve the truth value, meaning, or difficulty level of the original propositions. The prompt asks the model to paraphrase 'without changing the meaning,' yet no human evaluation, automated similarity check, or consistency re-test is described. If some paraphrases subtly alter proposition content, comparisons such as 'paraphrasing the correct options reduces the F1 score for Llama and Mistral' could reflect paraphrase artifacts rather than reasoning behavior. The authors should provide at least a sample-based human or model-based validation that paraphrases preserve factual content and difficulty.
- [Section 3.3, Phase 4] The dataset's ground-truth correctness labels and difficulty assignments rest primarily on a single CBE expert. The second, non-expert evaluator agreed with the correctness labels at only 67% F1, and agreement on the medium and high difficulty levels was 15% and 42% F1 respectively. Since the difficulty manipulation drives most of the cross-group comparisons in Groups 2-5, the low inter-evaluator agreement on Levels 2 and 3 is a real reliability concern. The authors should either recruit additional domain experts to validate the difficulty ordering, or restrict strong conclusions to difficulty levels where agreement is acceptable.
- [Table 1 and Section 4] The statistical support for several comparative claims is thin. The median F1 and standard deviation in Table 1 are computed over only four option-count configurations (5, 7, 10, and 15), and the text draws conclusions such as 'paraphrasing the correct options reduces the F1 score for Llama across all difficulty levels' from differences as small as 1.5-3.5 points. The paper should report per-claim or per-item variation, provide error bars, and use paired statistical tests across the shared claims to determine whether the observed differences are systematic rather than noise. Several statements about model ordering (e.g., 'Llama performs best, followed by Mistral and Gemma') would also benefit from such testing.
minor comments (6)
- [Section 3.2] The bullet numbering of the hypotheses is inconsistent: the list has '2. Introduce linguistic perturbations' followed by '2. Apply common sense' and then '3.' style numbering; the second item should be numbered 3.
- [Section 4, 'Effect of the difficulty of incorrect options'] The text reads 'hampers the model performance for Llama and Mistra' but the model name should be 'Mistral'.
- [Table 1 caption] The caption says 'median F1 and standard deviation across experiments', but the table appears to aggregate over the four option-count configurations per experiment. Please specify the aggregation unit and the number of examples per cell so readers can interpret the variance correctly.
- [Figure 4] The text in Section 4 states that 'Figure 4 plots the precision, recall, and F1 scores for Group 1 experiments,' but the caption of Figure 4 reads 'for experiments in comparison suite 5.' This mismatch should be corrected.
- [Limitations] The limitation statement says that 'extraneous experiments with larger and open-API models indicate a similar trend, but they are not quantified and non-generalizable.' This unquantified claim is not verifiable; either report the numbers or remove the statement.
- [Section 3.3, Phase 3] The sentence 'The CBE expert evaluated the clustering results, which were accurate in only 28% of cases' is ambiguous: it is unclear whether '28% of cases' refers to the proportion of clusters, proposition pairs, or individual propositions. Please clarify the denominator.
Circularity Check
No circularity: the benchmark is expert-curated and the evaluated models are external; no fitted parameter is renamed as a prediction.
full rationale
This paper is an empirical benchmark study rather than a derivation, so the main circularity patterns do not apply. The dataset construction uses LLM assistance, but the final content is expert-curated: Mistral-7B proposes propositions, yet a CBE expert edits, deletes, and adds items, and the standardization step is explicitly corrected because LLM clustering was only 28% accurate. The test set labels are manually constructed by the CBE expert and evaluated by a non-expert, with the acknowledged 67% F1 agreement indicating human labeling, not model-derived labels. The three evaluated LLMs are externally pretrained models; no parameter is fit to the benchmark and then reported as a prediction. The Group 1 'none of the above' inference is an empirical and construct-validity claim, not a tautology: the authors hypothesize that knowledgeable agents should choose 'none', observe a performance collapse, and infer limited reasoning. The paper itself flags that the position of the 'none' option could be a confounding variable, which is a legitimate threat to the conclusion but is not a circular reduction of the conclusion to its inputs. The only mild self-reference is using Llama-3.1-8B to generate paraphrases and then later evaluating Llama-3.1-8B on those paraphrased options; this is a tooling artifact that could affect results empirically, but it does not render the evaluation equivalent to its construction by definition. There is no load-bearing self-citation chain. Overall, the derivation chain is self-contained as an empirical benchmark, and no circular step can be exhibited from the paper's text.
Assumptions & free parameters
assumptions (4)
- domain assumption The CBE expert's proposition labels and entailment annotations are ground truth.
- domain assumption Textual entailment of propositions over a claim is a valid operationalization of p-knowledge (applied knowledge) for LLM reasoning in CBE.
- ad hoc to paper Llama-3.1-8B generated paraphrases preserve the truth value and meaning of the original propositions.
- ad hoc to paper Selecting 'none' when all options are false is a valid test of knowledgeability, unconfounded by model bias toward selecting options.
Cite this review
Pith. "Pith review of Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research." pith.science (2026). https://pith.science/paper/IVIV6KHB
@misc{pith2026250506964,
author = {Pith},
title = {Pith review of: Bridging AI and Carbon Capture: A Dataset for LLMs in Ionic Liquids and CBE Research},
year = {2026},
howpublished = {\url{https://pith.science/paper/IVIV6KHB}},
note = {Machine review of arXiv:2505.06964}
}
read the original abstract
Large Language Models (LLMs) have demonstrated exceptional performance in general knowledge and reasoning tasks across various domains. However, their effectiveness in specialized scientific fields like Chemical and Biological Engineering (CBE) remains underexplored. Addressing this gap requires robust evaluation benchmarks that assess both knowledge and reasoning capabilities in these niche areas, which are currently lacking. To bridge this divide, we present a comprehensive empirical analysis of LLM reasoning capabilities in CBE, with a focus on Ionic Liquids (ILs) for carbon sequestration - an emerging solution for mitigating global warming. We develop and release an expert - curated dataset of 5,920 examples designed to benchmark LLMs' reasoning in this domain. The dataset incorporates varying levels of difficulty, balancing linguistic complexity and domain-specific knowledge. Using this dataset, we evaluate three open-source LLMs with fewer than 10 billion parameters. Our findings reveal that while smaller general-purpose LLMs exhibit basic knowledge of ILs, they lack the specialized reasoning skills necessary for advanced applications. Building on these results, we discuss strategies to enhance the utility of LLMs for carbon capture research, particularly using ILs. Given the significant carbon footprint of LLMs, aligning their development with IL research presents a unique opportunity to foster mutual progress in both fields and advance global efforts toward achieving carbon neutrality by 2050.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Mahsa Aghaie, Nima Rezaei, and Sohrab Zendehboudi. 2018. A systematic review on co2 capture with ionic liquids: Current status and future prospects. Renewable and sustainable energy reviews, 96:502--525
2018
-
[2]
Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024. https://doi.org/10.18653/v1/2024.acl-long.671 Investigating cultural alignment of large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 12404--12422, Bangkok, Thailand. Association for Compu...
-
[3]
Jennifer L Anthony, Edward J Maginn, and Joan F Brennecke. 2002. Solubilities and thermodynamic properties of gases in the ionic liquid 1-n-butyl-3-methylimidazolium hexafluorophosphate. The Journal of Physical Chemistry B, 106(29):7315--7320
2002
-
[4]
John L Austin. 1961. Other minds
1961
-
[5]
Razvan Azamfirei, Sapna R Kudchadkar, and James Fackler. 2023. Large language models and the perils of their hallucinations. Critical Care, 27(1):120
2023
-
[6]
Igor Baskin, Alon Epshtein, and Yair Ein-Eli. 2022. Benchmarking machine learning methods for modeling physical properties of ionic liquids. Journal of Molecular Liquids, 351:118616
2022
-
[7]
Iz Beltagy, Kyle Lo, and Arman Cohan. 2019. Scibert: A pretrained language model for scientific text. arXiv preprint arXiv:1903.10676
arXiv 2019
-
[8]
Philippe Besnard and Anthony Hunter. 2008. Elements of argumentation, volume 47. MIT press Cambridge
2008
Show all 111 references
-
[9]
Lloyd F Bitzer. 2020. Aristotle's enthymeme revisited. In Landmark Essays on Aristotelian Rhetoric, pages 179--191. Routledge
2020
-
[10]
Lynnette A Blanchard, Zhiyong Gu, and Joan F Brennecke. 2001. High-pressure phase behavior of ionic liquid/co2 systems. The Journal of Physical Chemistry B, 105(12):2437--2444
2001
-
[11]
Lynnette A Blanchard, Dan Hancu, Eric J Beckman, and Joan F Brennecke. 1999. Green processing using ionic liquids and co2. Nature, 399(6731):28--29
1999
-
[12]
Andres M Bran, Sam Cox, Oliver Schilter, Carlo Baldassari, Andrew D White, and Philippe Schwaller. 2023. Chemcrow: Augmenting large-language models with chemistry tools. arXiv preprint arXiv:2304.05376
2023 arXiv
-
[13]
Tom B Brown. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165
2020 arXiv
-
[14]
S \'e bastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, et al. 2023. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712
2023 arXiv
-
[15]
Markus J Buehler. 2023 a . Generative pretrained autoregressive transformer graph neural network applied to the analysis and discovery of novel proteins. Journal of Applied Physics, 134(8)
2023
-
[16]
Markus J Buehler. 2023 b . Melm, a generative pretrained language modeling framework that solves forward and inverse mechanics problems. Journal of the Mechanics and Physics of Solids, 181:105454
2023
-
[17]
Lingdi Cao, Peng Zhu, Yongsheng Zhao, and Jihong Zhao. 2018. Using machine learning and quantum chemistry descriptors to predict the toxicity of ionic liquids. Journal of hazardous materials, 352:17--26
2018
-
[18]
Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023. https://doi.org/10.18653/v1/2023.c3nlp-1.7 Assessing cross-cultural alignment between C hat GPT and human societies: An empirical study . In Proceedings of the First Workshop on Cross-Cultur...
2023 doi
-
[19]
Cayque Monteiro Castro Nascimento and Andr \'e Silva Pimentel. 2023. Do large language models understand chemistry? a conversation with chatgpt. Journal of Chemical Information and Modeling, 63(6):1649--1655
2023
-
[20]
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240):1--113
2023
-
[21]
Le, Sergey Levine, and Yi Ma
Tianzhe Chu, Yuexiang Zhai, Jihan Yang, Shengbang Tong, Saining Xie, Dale Schuurmans, Quoc V. Le, Sergey Levine, and Yi Ma. 2025. https://arxiv.org/abs/2501.17161 Sft memorizes, rl generalizes: A comparative study of foundation model post-training . Preprint, arXiv:2501.17161
2025 arXiv
-
[22]
Pratik Dhakal and Jindal K Shah. 2022. A generalized machine learning model for predicting ionic conductivity of ionic liquids. Molecular Systems Design & Engineering, 7(10):1344--1353
2022
-
[23]
Radoslav S Dimitrov. 2016. The paris agreement on climate change: Behind closed doors. Global environmental politics, 16(3):1--11
2016
-
[24]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zha...
2024 arXiv
-
[25]
Ahmad Faiz, Sotaro Kaneda, Ruhan Wang, Rita Osi, Prateek Sharma, Fan Chen, and Lei Jiang. 2024. https://arxiv.org/abs/2309.14393 Llmcarbon: Modeling the end-to-end carbon footprint of large language models . arXiv preprint arXiv:2309.14393
2024 arXiv
-
[26]
Haijun Feng, Pingan Zhang, Wen Qin, Weiming Wang, and Huijing Wang. 2022. Estimation of solubility of acid gases in ionic liquids using different machine learning methods. Journal of Molecular Liquids, 349:118413
2022
-
[27]
Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau, and Anders S gaard. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.900 Defining knowledge: Bridging epistemology and large language models . In Proceedings of the 2024 Conference on Empirical Methods in Na...
2024 doi
-
[28]
Daan Frenkel and Berend Smit. 2023. Understanding molecular simulation: from algorithms to applications. Elsevier
2023
-
[29]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Meng Wang, and Haofen Wang. 2024. https://arxiv.org/abs/2312.10997 Retrieval-augmented generation for large language models: A survey . Preprint, arXiv:2312.10997
2024 arXiv
-
[30]
Yingqiang Ge, Wenyue Hua, Kai Mei, Juntao Tan, Shuyuan Xu, Zelong Li, Yongfeng Zhang, et al. 2024. Openagi: When llm meets domain experts. Advances in Neural Information Processing Systems, 36
2024
-
[31]
Jiahui Geng, Fengyu Cai, Yuxia Wang, Heinz Koeppl, Preslav Nakov, and Iryna Gurevych. 2023. A survey of confidence estimation and calibration in large language models. arXiv preprint arXiv:2311.08298
2023 arXiv
-
[32]
Joel Guiot and Wolfgang Cramer. 2016. Climate change: The 2015 paris agreement thresholds and mediterranean basin ecosystems. Science, 354(6311):465--468
2016
-
[33]
Taicheng Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh Chawla, Olaf Wiest, Xiangliang Zhang, et al. 2023. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Advances in Neural Information Processing Systems, 36:59662--59688
2023
-
[34]
Stefan Harrer. 2023. Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine. EBioMedicine, 90
2023
-
[35]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. https://arxiv.org/abs/2106.09685 Lora: Low-rank adaptation of large language models . Preprint, arXiv:2106.09685
2021 arXiv
-
[36]
Yiwen Hu and Markus J Buehler. 2022. End-to-end protein normal mode frequency predictions using language and graph models and application to sonification. ACS nano, 16(12):20656--20670
2022
-
[37]
Yiwen Hu and Markus J Buehler. 2023. Deep language models for interpretative and predictive materials science. APL Machine Learning, 1(1)
2023
-
[38]
Hsiu-Yuan Huang, Yutong Yang, Zhaoxi Zhang, Sanwoo Lee, and Yunfang Wu. 2024. A survey of uncertainty estimation in llms: Theory meets practice. arXiv preprint arXiv:2410.15326
2024 arXiv
-
[39]
Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. 2023. Look before you leap: An exploratory study of uncertainty measurement for large language models. arXiv preprint arXiv:2307.10236
2023 arXiv
-
[40]
Pascale Husson-Borg, Vladimir Majer, and Margarida F Costa Gomes. 2003. Solubilities of oxygen and carbon dioxide in butyl methyl imidazolium tetrafluoroborate as a function of temperature and at pressures close to atmospheric pressure. Journal of Chemical & Engineering Data, ...
2003
-
[41]
Pavan Inguva, Vijesh J Bhute, Thomas NH Cheng, and Pierre J Walker. 2021. Introducing students to research codes: A short course on solving partial differential equations in python. Education for Chemical Engineers, 36:1--11
2021
-
[42]
Kevin Maik Jablonka, Qianxiang Ai, Alexander Al-Feghali, Shruti Badhwar, Joshua D Bocarsly, Andres M Bran, Stefan Bringuier, L Catherine Brinson, Kamal Choudhary, Defne Circi, et al. 2023. 14 examples of how llms can transform materials science and chemistry: a reflection on a...
2023
-
[43]
Akshita Jha, Aida Mostafazadeh Davani, Chandan K Reddy, Shachi Dave, Vinodkumar Prabhakaran, and Sunipa Dev. 2023. https://doi.org/10.18653/v1/2023.acl-long.548 S ee GULL : A stereotype benchmark with broad geo-cultural coverage leveraging generative models . In Proceedings of...
2023 doi
-
[44]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1--38
2023
-
[45]
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...
2023 arXiv
-
[46]
John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, Kathryn Tunyasuvunakool, Russ Bates, Augustin Z \' dek, Anna Potapenko, et al. 2021. Highly accurate protein structure prediction with alphafold. nature, 596(7873):583--589
2021
-
[47]
Nikhil Kandpal, Haikang Deng, Adam Roberts, Eric Wallace, and Colin Raffel. 2023. Large language models struggle to learn long-tail knowledge. arXiv preprint arXiv:2211.08411
2023 arXiv
-
[48]
Eesha Khare, Constancio Gonzalez-Obeso, David L Kaplan, and Markus J Buehler. 2022. Collagentransformer: end-to-end transformer model to predict thermal stability of collagen triple helices using an nlp approach. ACS Biomaterials Science & Engineering, 8(10):4301--4310
2022
-
[49]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Proc...
2020
-
[50]
Huihan Li, Liwei Jiang, Nouha Dziri, Xiang Ren, and Yejin Choi. 2024 a . https://openreview.net/forum?id=DbsLm2KAqP CULTURE - GEN : Revealing global cultural perception in language models through natural language prompting . In First Conference on Language Modeling
2024
-
[51]
Huihan Li, Liwei Jiang, Jena D Hwang, Hyunwoo Kim, Sebastin Santy, Taylor Sorensen, Bill Yuchen Lin, Nouha Dziri, Xiang Ren, and Yejin Choi. 2024 b . Culture-gen: Revealing global cultural perception in language models through natural language prompting. arXiv preprint arXiv:2...
2024 arXiv
-
[52]
thirsty
Pengfei Li, Jianyi Yang, Mohammad A Islam, and Shaolei Ren. 2023. Making ai less" thirsty": Uncovering and addressing the secret water footprint of ai models. arXiv preprint arXiv:2304.03271
2023 arXiv
-
[53]
Weng Marc Lim, Asanka Gunasekara, Jessica Leigh Pallant, Jason Ian Pallant, and Ekaterina Pechenkina. 2023. Generative ai and the future of education: Ragnar \"o k or reformation? a paradoxical perspective from management educators. The international journal of management educ...
2023
-
[54]
Frank YC Liu, Bo Ni, and Markus J Buehler. 2022. Presto: Rapid protein mechanical strength prediction with an end-to-end deep learning model. Extreme Mechanics Letters, 55:101803
2022
-
[55]
Wei Lu, David L Kaplan, and Markus J Buehler. 2024. Generative modeling, design, and analysis of spider silk protein sequences for enhanced mechanical properties. Advanced Functional Materials, 34(11):2311324
2024
-
[56]
Wei Lu, Nic A Lee, and Markus J Buehler. 2023. Modeling and design of heterogeneous hierarchical bioinspired spider web structures using deep learning and additive manufacturing. Proceedings of the National Academy of Sciences, 120(31):e2305273120
2023
-
[57]
Rachel K Luu and Markus J Buehler. 2024. Bioinspiredllm: Conversational large language model for the mechanics of biological and bio-inspired materials. Advanced Science, 11(10):2306724
2024
-
[58]
Rachel K Luu, Marcin Wysokowski, and Markus J Buehler. 2023. Generative discovery of de novo chemical designs using diffusion modeling and transformer deep neural networks with application to deep eutectic solvents. Applied Physics Letters, 122(23)
2023
-
[59]
Edward J Maginn. 2009. Molecular simulation of ionic liquids: current status and future opportunities. Journal of Physics: Condensed Matter, 21(37):373101
2009
-
[60]
Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. 2022. Peft: State-of-the-art parameter-efficient fine-tuning methods. https://github.com/huggingface/peft
2022
-
[61]
Nick McKenna, Tianyi Li, Liang Cheng, Mohammad Javad Hosseini, Mark Johnson, and Mark Steedman. 2023. Sources of hallucination by large language models on inference tasks. arXiv preprint arXiv:2305.14552
2023 arXiv
-
[62]
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, Edouard Grave, Yann LeCun, and Thomas Scialom. 2023. Augmented language models: a survey. arXiv preprint arX...
2023 arXiv
-
[63]
Silvia Milano, Joshua A McGrane, and Sabina Leonelli. 2023. Large language models challenge the future of higher education. Nature Machine Intelligence, 5(4):333--334
2023
-
[64]
Kusuri Murakumo, Naruki Yoshikawa, Kentaro Rikimaru, Shogo Nakamura, Kairi Furui, Takamasa Suzuki, Hiroyuki Yamasaki, Yuki Nishigaya, Yuzo Takagi, and Masahito Ohue. 2023. Llm drug discovery challenge: A contest as a feasibility study on the utilization of large language model...
2023
-
[65]
Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. https://doi.org/10.18653/v1/2021.acl-long.416 S tereo S et: Measuring stereotypical bias in pretrained language models . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th Inte...
2021 doi
-
[66]
Rahul Nadkarni, David Wadden, Iz Beltagy, Noah A Smith, Hannaneh Hajishirzi, and Tom Hope. 2021. Scientific language models for biomedical knowledge base completion: an empirical study. arXiv preprint arXiv:2106.09700
2021 arXiv
-
[67]
Nikita Nangia, Clara Vania, Rasika Bhalerao, and Samuel R. Bowman. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.154 C row S -pairs: A challenge dataset for measuring social biases in masked language models . In Proceedings of the 2020 Conference on Empirical Methods in Na...
2020 doi
-
[68]
Robert Nozick. 2016. Knowledge and scepticism. In Readings in Formal Epistemology: Sourcebook, pages 587--603. Springer
2016
-
[69]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[70]
Kamil Paduszynski. 2016. In silico calculation of infinite dilution activity coefficients of molecular solutes in ionic liquids: critical review of current methods and new models based on three machine learning algorithms. Journal of chemical information and modeling, 56(8):1420--1437
2016
-
[71]
Saurabh Kumar Pandey, Harshit Budhiraja, Sougata Saha, and Monojit Choudhury. 2025. https://aclanthology.org/2025.coling-demos.21/ CULTURALLY YOURS : A reading assistant for cross-cultural content . In Proceedings of the 31st International Conference on Computational Linguisti...
2025
-
[72]
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. 2021. Carbon emissions and large neural network training. arXiv preprint arXiv:2104.10350
2021 arXiv
-
[73]
\'A lvaro P \'e rez-Salado Kamps, Dirk Tuma, Jianzhong Xia, and Gerd Maurer. 2003. Solubility of co2 in the ionic liquid [bmim][pf6]. Journal of Chemical & Engineering Data, 48(3):746--749
2003
-
[74]
Mahinder Ramdin, Theo W de Loos, and Thijs JH Vlugt. 2012. State-of-the-art of co2 capture with ionic liquids. Industrial & Engineering Chemistry Research, 51(24):8149--8177
2012
-
[75]
Abhinav Sukumar Rao, Aditi Khandelwal, Kumar Tanmay, Utkarsh Agarwal, and Monojit Choudhury. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.892 Ethical reasoning over moral alignment: A case and framework for in-context ethical policies in LLM s . In Findings of the Ass...
2023 doi
-
[76]
Nils Reimers and Iryna Gurevych. 2019. https://arxiv.org/abs/1908.10084 Sentence-bert: Sentence embeddings using siamese bert-networks . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics
2019 arXiv
-
[77]
Christopher J. Rhodes. 2016. https://doi.org/10.3184/003685016X14528569315192 The 2015 paris climate change conference: Cop21 . Science Progress, 99(1):97--104. PMID: 27120818
2016 doi
-
[78]
Matthias C Rillig, Marlene gerstrand, Mohan Bi, Kenneth A Gould, and Uli Sauerland. 2023. Risks and benefits of large language models for the environment. Environmental Science & Technology, 57(9):3464--3466
2023
-
[79]
Anthony Robbins. 2016. How to understand the results of the climate change summit: Conference of parties21 (cop21) paris 2015. Journal of public health policy, 37(2):129--132
2016
-
[80]
Sougata Saha, Saurabh Kumar Pandey, Harshit Gupta, and Monojit Choudhury. 2025. https://arxiv.org/abs/2502.09636 Reading between the lines: Can llms identify cross-cultural communication gaps? Preprint, arXiv:2502.09636
2025 arXiv
-
[81]
Eloy S Sanz-P \'e rez, Christopher R Murdock, Stephanie A Didas, and Christopher W Jones. 2016. Direct capture of co2 from ambient air. Chemical reviews, 116(19):11840--11876
2016
-
[82]
Crispin Sartwell. 1992. Why knowledge is merely true belief. The Journal of Philosophy, 89(4):167--180
1992
-
[83]
Timo Schick, Jane Dwivedi-Yu, Roberto Dess \` , Roberta Raileanu, Maria Lomeli, Eric Hambro, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. 2024. Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36
2024
-
[84]
Quintin R Sheridan, William F Schneider, and Edward J Maginn. 2018. Role of molecular modeling in the development of co2--reactive ionic liquids. Chemical reviews, 118(10):5242--5260
2018
-
[85]
Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in nlp. arXiv preprint arXiv:1906.02243
2019 arXiv
-
[86]
Kumar Tanmay, Aditi Khandelwal, Utkarsh Agarwal, and Monojit Choudhury. 2023. https://arxiv.org/abs/2309.13356 Probing the moral development of large language models through defining issues test . Preprint, arXiv:2309.13356
2023 arXiv
-
[87]
Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science. arXiv preprint arXiv:2211.09085
2022 arXiv
-
[88]
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan...
2024 arXiv
-
[89]
van Gunsteren, Xavier Daura, Niels Hansen, Alan E
Wilfred F. van Gunsteren, Xavier Daura, Niels Hansen, Alan E. Mark, Chris Oostenbrink, Sereina Riniker, and Lorna J. Smith. 2018. https://doi.org/10.1002/anie.201702945 Validation of molecular simulation: An overview of issues . Angewandte Chemie International Edition, 57(4):884--902
2018 doi
-
[90]
van Gunsteren and Alan E
Wilfred F. van Gunsteren and Alan E. Mark. 1998. https://doi.org/10.1063/1.476021 Validation of molecular dynamics simulation . The Journal of Chemical Physics, 108(15):6109--6116
1998 doi
-
[91]
Neeraj Varshney, Wenlin Yao, Hongming Zhang, Jianshu Chen, and Dong Yu. 2023. A stitch in time saves nine: Detecting and mitigating hallucinations of llms by actively validating low-confidence generation. arXiv preprint arXiv:2307.03987
2023 arXiv
-
[92]
Pablo Villalobos, Anson Ho, Jaime Sevilla, Tamay Besiroglu, Lennart Heim, and Marius Hobbhahn. 2024. Position: will we run out of data? limits of llm scaling based on human-generated data. In Proceedings of the 41st International Conference on Machine Learning, ICML'24. JMLR.org
2024
-
[93]
Pablo Villalobos, Jaime Sevilla, Lennart Heim, Tamay Besiroglu, Marius Hobbhahn, and Anson Ho. 2022. Will we run out of data? an analysis of the limits of scaling datasets in machine learning. arXiv preprint arXiv:2211.04325, 1
2022 arXiv
-
[94]
Douglas Walton, Christopher Reed, and Fabrizio Macagno. 2008. Argumentation schemes. Cambridge University Press
2008
-
[95]
Douglas N Walton. 1996. Argument structure: A pragmatic theory. University of Toronto Press Toronto
1996
-
[96]
Yixin Wan, Jieyu Zhao, Aman Chadha, Nanyun Peng, and Kai-Wei Chang. 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.648 Are personalized stochastic parrots more dangerous? evaluating persona biases in dialogue systems . In Findings of the Association for Computational Li...
2023 doi
-
[97]
Shaofei Wang, Xueqin Li, Hong Wu, Zhizhang Tian, Qingping Xin, Guangwei He, Dongdong Peng, Silu Chen, Yan Yin, Zhongyi Jiang, et al. 2016. Advances in high permeability polymer-based membrane materials for co2 separations. Energy & Environmental Science, 9(6):1863--1890
2016
-
[98]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou. 2023. Chain-of-thought prompting elicits reasoning in large language models. arXiv preprint arXiv:2201.11903
2023 arXiv
-
[99]
Andrew D White. 2023. The future of chemistry is language. Nature Reviews Chemistry, 7(7):457--458
2023
-
[100]
Timothy Williamson. 2005. Knowledge, context, and the agent’s. Contextualism in philosophy, page 91
2005
-
[101]
Fanghua Ye, Mingming Yang, Jianhui Pang, Longyue Wang, Derek F Wong, Emine Yilmaz, Shuming Shi, and Zhaopeng Tu. 2024. Benchmarking llms via uncertainty quantification. arXiv preprint arXiv:2401.12794
2024 arXiv
-
[102]
Chi-Hua Yu, Wei Chen, Yu-Hsuan Chiang, Kai Guo, Zaira Martin Moldes, David L Kaplan, and Markus J Buehler. 2022 a . End-to-end deep learning model to predict and design secondary structure content of structural proteins. ACS biomaterials science & engineering, 8(3):1156--1165
2022
-
[103]
Chi-Hua Yu, Eesha Khare, Om Prakash Narayan, Rachael Parker, David L Kaplan, and Markus J Buehler. 2022 b . Colgen: An end-to-end deep learning model to predict thermal stability of de novo collagen sequences. Journal of the mechanical behavior of biomedical materials, 125:104921
2022
-
[104]
Linda Zagzebski. 2017. What is knowledge? The Blackwell guide to epistemology, pages 92--116
2017
-
[105]
Stefano E Zanco, Jos \'e -Francisco P \'e rez-Calvo, Antonio Gas \'o s, Beatrice Cordiano, Viola Becattini, and Marco Mazzotti. 2021. Postcombustion co2 capture: a comparative techno-economic assessment of three technologies using a solvent, an adsorbent, and a membrane. ACS E...
2021
-
[106]
Shaojuan Zeng, Xiangping Zhang, Lu Bai, Xiaochun Zhang, Hui Wang, Jianji Wang, Di Bao, Mengdie Li, Xinyan Liu, and Suojiang Zhang. 2017. Ionic-liquid-based co2 capture systems: structure, interaction and process. Chemical reviews, 117(14):9625--9673
2017
-
[107]
Di Zhang, Wei Liu, Qian Tan, Jingdan Chen, Hang Yan, Yuliang Yan, Jiatong Li, Weiran Huang, Xiangyu Yue, Wanli Ouyang, et al. 2024. Chemllm: A chemical large language model. arXiv preprint arXiv:2402.06852
2024 arXiv
-
[108]
Haochen Zhao, Xiangru Tang, Ziran Yang, Xiao Han, Xuanzhi Feng, Yueqing Fan, Senhao Cheng, Di Jin, Yilun Zhao, Arman Cohan, et al. 2024. Chemsafetybench: Benchmarking llm safety on chemistry domain. arXiv preprint arXiv:2411.16736
2024 arXiv
-
[109]
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023. A survey of large language models. arXiv preprint arXiv:2303.18223
2023 arXiv
-
[110]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[111]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.