REVIEW 5 major objections 5 minor 118 references
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
T0 review · 5 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The TrustNLP workshop's six editions show a field moving from post-hoc interpretability to mechanistic verification and proactive control, with each capability launch triggering a lagged shift in trust research.
desk verdict A genuinely useful longitudinal map of TrustNLP with a new 144-paper dataset, but the numbers are sloppy and the cross-venue generalization is currently asserted, not demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a six-category trust taxonomy built from TrustLLM's dimensions, refined with DecodingTrust's finer-grained distinctions, plus Explainability added because the corpus demanded it. Papers are multi-labeled by three annotators—one human and two large language models following identical annotation instructions—with human adjudication on disagreements, yielding overall decision-level accuracy above 90 percent. This taxonomy is applied twice: once to all 144 TrustNLP papers and once, after a 30-keyword filter, to about 2,169 trust papers from four major NLP conferences, so the trends can be checked against a field-level baseline. The chronological synthesis organizes the same corpus into five phases—interpretability and bias, the generative pivot, trust as a trade-off problem, agentic and multimodal frontiers, and mechanistic trust and safety at scale—that map shifts in the dominant human–model interaction mode.
What would settle it
Classify trust-related papers from a different set of high-profile AI venues outside the NLP-conference family over 2021–2026 with the same six-dimension scheme and compare the year-by-year proportions; if truthfulness does not rise from absent to a dominant share after late 2022, or if explainability fails to rebound through mechanistic methods around 2026, the claimed field-wide transition and its lag structure would not generalize beyond TrustNLP.
Extended reading notes
Core claim
The central claim is that TrustNLP's six-year record documents a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems, and that this transition was driven by capability emergence rather than by cumulative theoretical progress. The paper's quantitative evidence is a classification of all 144 archival proceedings papers into six trust dimensions—Fairness & Bias, Robustness & Adversarial, Privacy, Machine Ethics & Safety, Truthfulness, and Explainability—derived from TrustLLM and DecodingTrust. The counts show truthfulness appearing only after the generative pivot and reaching 38 total papers, fairness present in every edition with 30 total, and explainability falling from 9 papers in 2021–2022 to 2 in 2025 before rebounding to 13 in 2026 through mechanistic interpretability. The authors read these trends as a lagged response to capability events: chat models activated all dimensions at once in 2023, open-weight and frontier models brought trade-off research in 2024, and agentic and multimodal systems pushed safety alignment and mechanistic verification in 2025–2026. They further argue from a cross-venue comparison of about 2,000 papers that TrustNLP's distribution mirrors the broader NLP community's priorities.
Load-bearing premise
The load-bearing premise is that TrustNLP's six editions faithfully represent the trust research priorities of the entire NLP community; if the workshop's organizers, submission pool, or venue placement attract a skewed slice of the field, the claimed field-wide transition is really only a workshop-level trend.
Editorial extensions
If this is right
- The truthfulness surge means hallucination, factuality, and calibration are now leading trust problems across the field, not niche concerns.
- Mechanistic interpretability—probing and editing model internals—will keep displacing post-hoc attribution as the standard way to answer why a model behaved as it did.
- Past capability shocks each produced a lagged shift in trust topics, so the next major modeling change should be followed within about a year by a visible jump in corresponding trust papers.
- Output-level evaluation alone will increasingly be seen as insufficient, making internal probes a first-class component of trustworthy-evaluation practice.
- Without a unifying framework, each new capability will continue to generate disconnected trust research lines rather than cumulative progress.
Reading between the lines
- A testable extension would be to apply the same six-dimension classification to the next two workshop editions and check whether the observed one-edition lag between capability releases and topical shifts persists, since the authors flag that the regularity is retrospective.
- The limitations section implies a sharper check the authors did not run: classifying trust papers at non-NLP-specific AI venues over the same period would reveal whether robustness or privacy grows faster than truthfulness there, which would narrow the representativeness claim to NLP venues.
- An implicit consequence of the outputs-are-not-internals insight is that regulatory or deployment audits relying only on behavioral tests will miss latent misalignment, so audit protocols may soon need representation-level checks as a standard component.
- The observed benchmark saturation suggests that evaluation sets should be maintained adversarially and updated as models improve; otherwise static benchmarks risk becoming training data and losing their measurement signal.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes all 144 archival TrustNLP workshop papers from 2021 to 2026, classifying them along six trust dimensions (Fairness & Bias, Robustness & Adversarial, Privacy, Machine Ethics & Safety, Truthfulness, Explainability) using one human and two LLM annotators. It reports temporal trends, a phase-by-phase synthesis of technical contributions, four structural insights, and a cross-venue comparison with ACL, NAACL, EACL, and EMNLP intended to show that TrustNLP mirrors the broader field's trust priorities. The abstract frames the main finding as a field-wide shift from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems.
Significance. If the claims hold, the paper would provide a valuable longitudinal record of how a stable ACL-affiliated workshop responded to capability changes in LLMs, and its full annotation table (Appendix D) is a useful resource for future meta-analyses. The use of two independent LLM classifiers with human adjudication, with agreement statistics reported, is a strength. However, the central generalization that TrustNLP is representative of the broader NLP community rests on Appendix C, and that evidence is currently not valid as implemented, due to inconsistent inclusion criteria and missing statistical support. The internal trend analysis from Table 2 is self-contained and useful even if the representativeness claim is withdrawn.
major comments (5)
- [Abstract, §3.3, Table 2] The truthfulness figures are internally inconsistent. The abstract states that truthfulness comprises 37% of papers by 2025–2026, but Table 2 gives 14 (2025) + 13 (2026) = 27 papers out of 79 for those two years, which is 34.2%, while the body's §3.3 reports 38 total and 26% of the corpus. The 37% figure appears to match only the single year 2025 (14/38 = 36.8%). Please reconcile these numbers, since the claim that truthfulness is the fastest-growing dimension is a headline quantitative result.
- [Figure 2, Appendix C] Figure 2 labels the TrustNLP sample as n=136, while the paper states it analyzes all 144 proceedings papers. The discrepancy of 8 papers is never explained. Either the figure is using a subset (for example, excluding papers that failed classification), or the corpus count is wrong. This must be clarified because the figure is the only quantitative bridge to the field-generalization claim.
- [Appendix C] The cross-venue comparison does not apply the same inclusion rule to TrustNLP and the *CL venues. The ACL/NAACL/EACL/EMNLP set is filtered by title keywords, while TrustNLP appears to be entered as the full workshop corpus, so the comparison conflates the effect of keyword filtering with the effect of venue selection. In addition, Appendix C reports 2,169 papers both as the pre-exclusion candidate set and as the post-exclusion classified set, and Figure 2 shows only aggregate proportions with no per-year breakdown, no error bars, and no statistical test. Because this appendix is the only quantitative support for the paper's central claim that TrustNLP mirrors the field, the comparison must be recomputed with identical filtering and accompanied by uncertainty quantification, or the claim must be narrowed.
- [§4.5, Table 2] There are numerical mismatches between the narrative and Table 2 for the 2026 edition: §4.5 says explainability research comprises 11 papers in 2026, while Table 2 reports 13 Explainability papers; §4.5 also says Machine Ethics & Safety reaches its highest count with 8 papers, while Table 2 lists 9. Please correct the prose or the table so that the trend claims are reproducible from the data.
- [§5.1 Insight 1, §7] The claim of a causal or quasi-causal 'reactive lag' between capability events and topic shifts is stated as a structural insight, but the paper provides no quantitative test of the lag, and §7 itself acknowledges that the regularity is observed across only four capability emergences and that its future validity is an empirical question. The language in Insight 1 and the conclusion should be softened to a documented correlation, or a statistical analysis of event-to-topic lags should be added.
minor comments (5)
- [Figure 1 caption] The caption contains a typo: 'T oronto' should be 'Toronto'.
- [References] The reference to 'Bui and V on Der Wense' contains stray spaces; the author's name should be 'Katharina von der Wense'.
- [§3.3] The description of explainability as following a 'U-shaped trajectory' is not well supported by Table 2, which shows 5, 4, 2, 3, 2, 13 across years. The 2026 value is a sharp single-year spike rather than a smooth U-shape, so the characterization should be qualified or supported with additional analysis.
- [Appendix D] Several entries in Table 3 use an em dash in the human or model columns (for example, the 2022 'An Encoder Attribution Analysis' row has human label '—'). The paper should state explicitly whether a dash means 'not a primary dimension' or 'missing annotation', since this affects how the final label set in Table 2 was derived.
- [§3.2] The sentence 'In this way, we expand the scope of each class by merging overlapping class labels from the two taxonomies' is vague; please give a concrete example of how a TrustLLM label and a DecodingTrust label were merged into one of the six final dimensions.
Circularity Check
No significant circularity; the analysis is a retrospective synthesis with one minor self-referential taxonomy choice.
-
self definitional
[Section 3.1, Classification Taxonomy (Explainability addition); Section 3.3, Trends (Explainability trajectory)]
"However, explainability dominates the TrustNLP corpus in 2021–2022 and appears explicitly in the workshop CfP, so we add it as a corpus-motivated dimension. ... Explainability peaked early (9 papers across 2021–2022), declined to 2 papers in 2025, but rebounded sharply to 13 in 2026."
The Explainability category is added to the six-dimension taxonomy because the authors observe that explainability dominates the 2021–2022 portion of the same corpus. The subsequent finding that explainability 'peaked early' in 2021–2022 therefore restates the inclusion criterion rather than being an independent discovery. However, the decline to 2 papers and the 2026 resurgence to 13 are not entailed by the inclusion rule, so the circularity is limited to the early-peak statement and does not affect the paper's central claims about truthfulness growth or the cross-venue comparison.
full rationale
The paper's main quantitative claims are derived from a transparent annotation process using an externally grounded taxonomy (TrustLLM, DecodingTrust), with labels checked against two independent LLM classifiers and human adjudication. The cross-venue comparison in Appendix C is methodologically imperfect — the *CL set is title-keyword filtered while TrustNLP is entered as the full workshop corpus, the reported 2,169 papers are unchanged after excluding API errors, and no statistical test accompanies Figure 2 — but these are validity concerns, not circularity: the comparative distributions are not forced by the definitions. The abstract's 37% truthfulness figure is arithmetically inconsistent with Table 2's totals, again a correctness issue rather than a circular reduction. The author-organizer overlap and use of self-edited proceedings as the data source are natural for a workshop retrospective and do not constitute a load-bearing self-citation chain. The only mildly self-referential element is the corpus-motivated addition of the Explainability dimension, which presupposes the early dominance that the trends section then reports; this is minor and does not undermine the independent content of the analysis. Overall score 1 reflects one minor self-referential taxonomy choice with no load-bearing circularity.
Assumptions & free parameters
assumptions (4)
- ad hoc to paper Trust taxonomy derived by merging TrustLLM and DecodingTrust with a corpus-motivated Explainability dimension
- domain assumption Papers are classified using only title and abstract
- domain assumption The human author annotation is treated as ground truth for adjudication
- ad hoc to paper The reactive lag between capability events and topic shifts is interpreted as causal
Cite this review
Pith. "Pith review of From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop." pith.science (2026). https://pith.science/paper/UWZIEJYK
@misc{pith2026260811171,
author = {Pith},
title = {Pith review of: From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop},
year = {2026},
howpublished = {\url{https://pith.science/paper/UWZIEJYK}},
note = {Machine review of arXiv:2608.11171}
}
read the original abstract
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proceedings papers to 41 over six editions, documenting a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems. We synthesize insights from all 144 proceedings papers, classifying them along six trust dimensions grounded in established frameworks (TrustLLM, DecodingTrust). We observe co-occurrences with capability emergence. The release of the first high-impact chat models activated all trust dimensions simultaneously, while subsequent model generations shifted focus toward truthfulness and safety alignment. Analysis from the classification study reveals that truthfulness is the fastest-growing dimension (absent in 2021-2022, comprising 37% of papers by 2025-2026), fairness remains the most consistent theme, and explainability exhibits a U-shaped trajectory; declining as post-hoc methods lost relevance but resurging in 2026 through mechanistic interpretability. A cross-venue comparison with ACL, NAACL, EACL, and EMNLP (~2K papers) in the same period shows that TrustNLP's topical distribution closely follows the field average. We identify four structural insights and conclude with actionable directions for the research community.
Figures
Reference graph
Works this paper leans on
-
[2]
Muhammad Adilazuarda. 2024. https://aclanthology.org/2024.trustnlp-1.1/ Beyond T uring: A comparative analysis of approaches for detecting machine-generated text . In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024), Mexico City, Mexico. Association for Computational Linguistics
2024
-
[6]
Anthropic. 2023. https://www-cdn.anthropic.com/bd2a28d2535bfb0494cc8e2a3bf135d2e7523226/Model-Card-Claude-2.pdf Model card and evaluations for claude models . Anthropic Model Card
2023
-
[10]
Esma Balkir, Svetlana Kiritchenko, Isar Nejadgholi, and Kathleen Fraser. 2022. https://aclanthology.org/2022.trustnlp-1.8/ Challenges in applying explainability methods to improve the fairness of NLP models . In Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022), pages 80--92, Seattle, U.S.A. Association for Computa...
2022
-
[11]
Stephanie Brandl, Emanuele Bugliarello, and Ilias Chalkidis. 2024. https://aclanthology.org/2024.trustnlp-1.10/ On the interplay between fairness and explainability . In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024), pages 94--108, Mexico City, Mexico. Association for Computational Linguistics
2024
-
[14]
Minh Duc Bui and Katharina Von Der Wense. 2024. https://aclanthology.org/2024.trustnlp-1.4/ The trade-off between performance, efficiency, and fairness in adapter modules for text classification . In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024), Mexico City, Mexico. Association for Computational Linguistics
2024
-
[15]
Trista Cao, Anubrata Das, Tharindu Kumarage, Yixin Wan, Satyapriya Krishna, Ninareh Mehrabi, Jwala Dhamala, Anil Ramakrishna, Aram Galystan, Anoop Kumar, Rahul Gupta, and Kai-Wei Chang, editors. 2025. https://aclanthology.org/2025.trustnlp-main.0/ Proceedings of the 5th Workshop on Trustworthy NLP ( TrustNLP 2025) . Association for Computational Linguisti...
2025
-
[26]
Emma Harvey, Allison Koenecke, and Rene F Kizilcec. 2025. https://doi.org/10.1145/3706598.3713210 " don't forget the teachers": Towards an educator-centered understanding of harms from large language models in education . In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1--19
arXiv 2025
-
[28]
Saghar Hosseini, Hamid Palangi, and Ahmed Hassan Awadallah. 2023. https://aclanthology.org/2023.trustnlp-1.11/ An empirical study of metrics to measure representational harms in pre-trained language models . In Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023), pages 121--134, Toronto, Canada. Association for Compu...
2023
Show all 118 references
-
[30]
Yue Huang, Lichao Sun, Haoran Wang, Siyuan Wu, Qihui Zhang, Yuan Li, Chujie Gao, Yixin Huang, Wenhan Lyu, Yixuan Zhang, and 1 others. 2024. https://arxiv.org/abs/2401.05561 T rust LLM : Trustworthiness in large language models . arXiv preprint arXiv:2401.05561
2024 arXiv
-
[31]
Sullam Jeoung, Jana Diesner, and Halil Kilicoglu. 2023. https://aclanthology.org/2023.trustnlp-1.7/ Examining the causal impact of first names on language models: The case of social commonsense reasoning . In Proceedings of the 3rd Workshop on Trustworthy Natural Language Proc...
2023
-
[33]
Heegyu Kim and Hyunsouk Cho. 2025. https://aclanthology.org/2025.trustnlp-main.7/ Break the breakout: Reinventing LM defense against jailbreak attacks with self-refine . In Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), Albuquerque, New Mexico. Association...
2025
-
[34]
Yash Kumar Lal, Preethi Lahoti, Aradhana Sinha, Yao Qin, and Ananth Balashankar. 2024. https://aclanthology.org/2024.trustnlp-1.2/ Automated adversarial discovery for safety classifiers . In Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2...
2024
-
[35]
Minghan Li, Xueguang Ma, and Jimmy Lin. 2022. https://aclanthology.org/2022.trustnlp-1.1/ An encoder attribution analysis for dense passage retriever in open-domain question answering . In Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 202...
2022
-
[43]
Anaelia Ovalle, Kai-Wei Chang, Yang Trista Cao, Ninareh Mehrabi, Jieyu Zhao, Aram Galstyan, Jwala Dhamala, Anoop Kumar, and Rahul Gupta, editors. 2024. https://aclanthology.org/2024.trustnlp-1.0/ Proceedings of the 4th Workshop on Trustworthy Natural Language Processing ( Trus...
2024
-
[44]
Anaelia Ovalle, Kai-Wei Chang, Ninareh Mehrabi, Yada Pruksachatkun, Aram Galystan, Jwala Dhamala, Apurv Verma, Trista Cao, Anoop Kumar, and Rahul Gupta, editors. 2023. https://aclanthology.org/2023.trustnlp-1.0/ Proceedings of the 3rd Workshop on Trustworthy Natural Language P...
2023
-
[47]
Yada Pruksachatkun, Anil Ramakrishna, Kai-Wei Chang, Satyapriya Krishna, Jwala Dhamala, Tanaya Guha, and Xiang Ren, editors. 2021. https://aclanthology.org/2021.trustnlp-1.0/ Proceedings of the First Workshop on Trustworthy Natural Language Processing . Association for Computa...
2021
-
[54]
Zhengyan Shi, Giuseppe Castellucci, Simone Filice, Saar Kuzi, Elad Kravi, Eugene Agichtein, Oleg Rokhlenko, and Shervin Malmasi. 2025. https://aclanthology.org/2025.trustnlp-main.4/ Ambiguity detection and uncertainty calibration for question answering with large language mode...
2025
-
[59]
Apurv Verma, Yada Pruksachatkun, Kai-Wei Chang, Aram Galstyan, Jwala Dhamala, and Yang Trista Cao, editors. 2022. https://aclanthology.org/2022.trustnlp-1.0/ Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing ( TrustNLP 2022) . Association for Computati...
2022
-
[61]
Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, and 1 others. 2023. https://arxiv.org/abs/2306.11698 DecodingTrust : A comprehensive assessment of trustworthiness in GPT models . In Advances i...
2023 arXiv
-
[64]
Kyra Yee, Alice Schoenauer Sebag, Olivia Redfield, Matthias Eck, Emily Sheng, and Luca Belli. 2023. https://aclanthology.org/2023.trustnlp-1.10/ A keyword based approach to understanding the overpenalization of marginalized groups by E nglish marginal abuse models on T witter ...
2023
-
[68]
Proceedings of the First Workshop on Trustworthy Natural Language Processing. 2021
2021
-
[69]
Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing ( TrustNLP 2022). 2022
2022
-
[70]
Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing ( TrustNLP 2023). 2023
2023
-
[71]
Proceedings of the 4th Workshop on Trustworthy Natural Language Processing ( TrustNLP 2024). 2024
2024
-
[72]
Proceedings of the 5th Workshop on Trustworthy NLP ( TrustNLP 2025). 2025
2025
-
[73]
2026 , howpublished =
Principled Design for Trustworthy. 2026 , howpublished =
2026
-
[74]
Human-Centered Explainable AI : Towards a Reflective Sociotechnical Approach
Ehsan, Upol and Riedl, Mark O. Human-Centered Explainable AI : Towards a Reflective Sociotechnical Approach. HCI International 2020 -- Late Breaking Papers: Multimodality and Intelligence. 2020. doi:10.1007/978-3-030-60117-1_33
2020 doi
-
[75]
and Wintersberger, Philipp and Manger, Carina and Hubig, Nina and Savage, Saiph and Weisz, Justin D
Ehsan, Upol and Watkins, Elizabeth A. and Wintersberger, Philipp and Manger, Carina and Hubig, Nina and Savage, Saiph and Weisz, Justin D. and Riener, Andreas , booktitle =. New Frontiers of Human-centered Explainable. 2025 , isbn =
2025
-
[76]
arXiv preprint arXiv:2406.03712 , year=
A survey on medical large language models: Technology, application, trustworthiness, and future directions , author=. arXiv preprint arXiv:2406.03712 , year=
-
[77]
arXiv preprint arXiv:2502.15865 , year=
Standard Benchmarks Fail--Auditing LLM Agents in Finance Must Prioritize Risk , author=. arXiv preprint arXiv:2502.15865 , year=
-
[78]
Don't Forget the Teachers
" Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education , author=. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems , pages=. 2025 , doi=
2025
-
[79]
2025 Silicon Valley Cybersecurity Conference (SVCC) , pages=
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies , author=. 2025 Silicon Valley Cybersecurity Conference (SVCC) , pages=. 2025 , organization=. doi:10.1109/sv...
2025
-
[80]
2023 , url =
Wang, Boxin and Chen, Weixin and Pei, Hengzhi and Xie, Chulin and Kang, Mintong and Zhang, Chenhui and Xu, Chejian and Xiong, Zidi and Dutta, Ritik and Schaeffer, Rylan and others , booktitle =. 2023 , url =
2023
-
[81]
2024 , url=
Huang, Yue and Sun, Lichao and Wang, Haoran and Wu, Siyuan and Zhang, Qihui and Li, Yuan and Gao, Chujie and Huang, Yixin and Lyu, Wenhan and Zhang, Yixuan and others , journal=. 2024 , url=
2024
-
[82]
Trustworthy LLM s: A Survey and Guideline for Evaluating Large Language Models' Alignment
Liu, Yang and Yao, Yuanshun and Ton, Jean-Francois and Zhang, Xiaoying and Guo, Ruocheng and Cheng, Hao and Klochkov, Yegor and Taufiq, Muhammad Faaiz and Li, Hang. Trustworthy LLM s: A Survey and Guideline for Evaluating Large Language Models' Alignment. arXiv preprint arXiv:...
-
[83]
2025 , publisher=
Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency , editor=. 2025 , publisher=
2025
-
[84]
Proceedings of the Eighth AAAI/ACM Conference on AI, Ethics, and Society (AIES-25) -- Main Track I , year =
-
[85]
Interpretability Rules: Jointly Bootstrapping a Neural Relation Extractor with an Explanation Decoder
Tang, Zheng and Surdeanu, Mihai. Interpretability Rules: Jointly Bootstrapping a Neural Relation Extractor with an Explanation Decoder. Proceedings of the First Workshop on Trustworthy Natural Language Processing. 2021. doi:10.18653/v1/2021.trustnlp-1.1
2021 doi
-
[86]
Measuring Biases of Word Embeddings: What Similarity Measures and Descriptive Statistics to Use?
Azarpanah, Hossein and Farhadloo, Mohsen. Measuring Biases of Word Embeddings: What Similarity Measures and Descriptive Statistics to Use?. Proceedings of the First Workshop on Trustworthy Natural Language Processing. 2021. doi:10.18653/v1/2021.trustnlp-1.2
2021 doi
-
[87]
and Kiritchenko, Svetlana and Balkir, Esma
Fraser, Kathleen C. and Kiritchenko, Svetlana and Balkir, Esma. Does Moral Code Have a Moral Code? Probing D elphi's Moral Philosophy. Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022). 2022. doi:10.18653/v1/2022.trustnlp-1.3
2022 doi
-
[88]
GPT s Don ' t Keep Secrets: Searching for Backdoor Watermark Triggers in Autoregressive Language Models
Lucas, Evan and Havens, Timothy. GPT s Don ' t Keep Secrets: Searching for Backdoor Watermark Triggers in Autoregressive Language Models. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.21
2023 doi
-
[89]
Reliability Check: An Analysis of GPT -3's Response to Sensitive Topics and Prompt Wording
Khatun, Aisha and Brown, Daniel. Reliability Check: An Analysis of GPT -3's Response to Sensitive Topics and Prompt Wording. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.8
2023 doi
-
[90]
Driving Context into Text-to-Text Privatization
Arnold, Stefan and Yesilbas, Dilara and Weinzierl, Sven. Driving Context into Text-to-Text Privatization. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.2
2023 doi
-
[91]
Expanding Scope: Adapting E nglish Adversarial Attacks to C hinese
Liu, Hanyu and Cai, Chengyuan and Qi, Yanjun. Expanding Scope: Adapting E nglish Adversarial Attacks to C hinese. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.24
2023 doi
-
[92]
Flatness-Aware Gradient Descent for Safe Conversational AI
Khalatbari, Leila and Hosseini, Saeid and Sameti, Hossein and Fung, Pascale. Flatness-Aware Gradient Descent for Safe Conversational AI. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024. doi:10.18653/v1/2024.trustnlp-1.15
2024 doi
-
[93]
PBI -Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
Cheng, Ruoxi and Ding, Yizhong and Cao, Shuirong and Duan, Ranjie and Jia, Xiaoshuang and Yuan, Shaowei and Wang, Zhiqiang and Jia, Xiaojun. PBI -Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization. Proceedings of the 5th Workshop on T...
2025 doi
-
[94]
Beyond Text-to- SQL for IoT Defense: A Comprehensive Framework for Querying and Classifying IoT Threats
Pavlich, Ryan and Ebadi, Nima and Tarbell, Richard and Linares, Billy and Tan, Adrian and Humphreys, Rachael and Das, Jayanta and Ghandiparsi, Rambod and Haley, Hannah and George, Jerris and Slavin, Rocky and Choo, Kim-Kwang Raymond and Dietrich, Glenn and Rios, Anthony. Beyon...
2025 doi
-
[95]
Minimal Evidence Group Identification for Claim Verification
Li, Xiangci and Chen, Sihao and Kapadia, Rajvi and Ouyang, Jessica and Zhang, Fan. Minimal Evidence Group Identification for Claim Verification. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653/v1/2025.trustnlp-main.8
2025 doi
-
[96]
Estimating Knowledge in Large Language Models Without Generating a Single Token
Gottesman, Daniela and Geva, Mor. Estimating Knowledge in Large Language Models Without Generating a Single Token. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. doi:10.18653/v1/2024.emnlp-main.232
2024 doi
-
[97]
Intrinsic Test of Unlearning Using Parametric Knowledge Traces
Hong, Yihuai and Yu, Lei and Yang, Haiqin and Ravfogel, Shauli and Geva, Mor. Intrinsic Test of Unlearning Using Parametric Knowledge Traces. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. doi:10.18653/v1/2025.emnlp-main.985
2025 doi
-
[98]
Can we trust the evaluation on C hat GPT ?
Aiyappa, Rachith and An, Jisun and Kwak, Haewoon and Ahn, Yong-Yeol. Can we trust the evaluation on C hat GPT ?. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.5
2023 doi
-
[99]
Improving Factuality of Abstractive Summarization via Contrastive Reward Learning
Chern, I-chun and Wang, Zhiruo and Das, Sanjan and Sharma, Bhavuk and Liu, Pengfei and Neubig, Graham. Improving Factuality of Abstractive Summarization via Contrastive Reward Learning. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023)....
2023 doi
-
[100]
Exploring Causal Mechanisms for Machine Text Detection Methods
Yoo, Kiyoon and Ahn, Wonhyuk and Song, Yeji and Kwak, Nojun. Exploring Causal Mechanisms for Machine Text Detection Methods. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024. doi:10.18653/v1/2024.trustnlp-1.7
2024 doi
-
[101]
On the Robustness of Agentic Function Calling
Rabinovich, Ella and Anaby Tavor, Ateret. On the Robustness of Agentic Function Calling. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653/v1/2025.trustnlp-main.20
2025 doi
-
[102]
Cross-Task Defense: Instruction-Tuning LLM s for Content Safety
Fu, Yu and Xiao, Wen and Chen, Jia and Li, Jiachen and Papalexakis, Evangelos and Chien, Aichi and Dong, Yue. Cross-Task Defense: Instruction-Tuning LLM s for Content Safety. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024. doi:...
2024 doi
-
[103]
Gender Bias in Natural Language Processing Across Human Languages
Matthews, Abigail and Grasso, Isabella and Mahoney, Christopher and Chen, Yan and Wali, Esma and Middleton, Thomas and Njie, Mariama and Matthews, Jeanna. Gender Bias in Natural Language Processing Across Human Languages. Proceedings of the First Workshop on Trustworthy Natura...
2021 doi
-
[104]
Into the Gap between What Language Models Say and What They Know
Geva, Mor. Into the Gap between What Language Models Say and What They Know. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[105]
The False Sense of Privacy in LLM s: Non-Verbatim Memorization and Semantic Leakage
Mireshghallah, Niloofar. The False Sense of Privacy in LLM s: Non-Verbatim Memorization and Semantic Leakage. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[106]
and Raimundo, Marcos M
Resck, Lucas E. and Raimundo, Marcos M. and Poco, Jorge. Exploring the Trade-off between Model Performance and Explanation Plausibility of Text Classifiers Using Human Rationales. Findings of the Association for Computational Linguistics: NAACL 2024. 2024
2024
-
[107]
2023 , howpublished =
Responsible. 2023 , howpublished =
2023
-
[108]
Nature Machine Intelligence , volume=
Common pitfalls and recommendations for using machine learning to detect and prognosticate for COVID-19 using chest radiographs and CT scans , author=. Nature Machine Intelligence , volume=. 2021 , month=. doi:10.1038/s42256-021-00307-0 , url=
2021 doi
-
[109]
Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI , year =
Jacovi, Alon and Marasovi\'. Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI , year =. Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency , pages =. doi:10.1145/3442188.3445923 , abstract =
2021
-
[110]
FAccT 2022 , year =
Sharaf, Amr and Daumé III, Hal and Ni, Renkun , title =. FAccT 2022 , year =
2022
-
[111]
SODAPOP : Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models
An, Haozhe and Li, Zongxia and Zhao, Jieyu and Rudinger, Rachel. SODAPOP : Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. doi:10.18...
2023 doi
-
[112]
F air B elief - Assessing Harmful Beliefs in Language Models
Setzu, Mattia and Marchiori Manerba, Marta and Minervini, Pasquale and Nozza, Debora. F air B elief - Assessing Harmful Beliefs in Language Models. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024. doi:10.18653/v1/2024.trustnlp-1.3
2024 doi
-
[113]
Investigating and Addressing Hallucinations of LLM s in Tasks Involving Negation
Varshney, Neeraj and Raj, Satyam and Mishra, Venkatesh and Chatterjee, Agneet and Saeidi, Amir and Sarkar, Ritika and Baral, Chitta. Investigating and Addressing Hallucinations of LLM s in Tasks Involving Negation. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2...
2025 doi
-
[114]
Introducing G en C eption for Multimodal LLM Benchmarking: You May Bypass Annotations
Cao, Lele and Buchner, Valentin and Senane, Zineb and Yang, Fangkai. Introducing G en C eption for Multimodal LLM Benchmarking: You May Bypass Annotations. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024. doi:10.18653/v1/2024.tr...
2024 doi
-
[115]
Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models
Zarharan, Majid and Wullschleger, Pascal and Behkam Kia, Babak and Pilehvar, Mohammad Taher and Foster, Jennifer. Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustN...
2024 doi
-
[116]
Disentangling Linguistic Features with Dimension-Wise Analysis of Vector Embeddings
Karwa, Saniya and Singh, Navpreet. Disentangling Linguistic Features with Dimension-Wise Analysis of Vector Embeddings. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653/v1/2025.trustnlp-main.30
2025 doi
-
[117]
On The Real-world Performance of Machine Translation: Exploring Social Media Post-authors' Perspectives
Gupta, Ananya and Takeuchi, Jae and Knijnenburg, Bart. On The Real-world Performance of Machine Translation: Exploring Social Media Post-authors' Perspectives. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/202...
2023 doi
-
[118]
V i B e: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
Rawte, Vipula and Jain, Sarthak and Sinha, Aarush and Kaushik, Garv and Bansal, Aman and Vishwanath, Prathiksha Rumale and Jain, Samyak Rajesh and Reganti, Aishwarya Naresh and Jain, Vinija and Chadha, Aman and Sheth, Amit and Das, Amitava. V i B e: A Text-to-Video Benchmark f...
2025
-
[119]
FACTOID : FAC tual en T ailment f O r halluc I nation Detection
Rawte, Vipula and Tonmoy, S.m Towhidul Islam and Nag, Shravani and Chadha, Aman and Sheth, Amit and Das, Amitava. FACTOID : FAC tual en T ailment f O r halluc I nation Detection. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653/v1/2025.tru...
2025 doi
-
[120]
Private Release of Text Embedding Vectors
Feyisetan, Oluwaseyi and Kasiviswanathan, Shiva. Private Release of Text Embedding Vectors. Proceedings of the First Workshop on Trustworthy Natural Language Processing. 2021. doi:10.18653/v1/2021.trustnlp-1.3
2021 doi
-
[121]
Challenges in Applying Explainability Methods to Improve the Fairness of NLP Models
Balkir, Esma and Kiritchenko, Svetlana and Nejadgholi, Isar and Fraser, Kathleen. Challenges in Applying Explainability Methods to Improve the Fairness of NLP Models. Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022). 2022
2022
-
[122]
An Encoder Attribution Analysis for Dense Passage Retriever in Open-Domain Question Answering
Li, Minghan and Ma, Xueguang and Lin, Jimmy. An Encoder Attribution Analysis for Dense Passage Retriever in Open-Domain Question Answering. Proceedings of the 2nd Workshop on Trustworthy Natural Language Processing (TrustNLP 2022). 2022
2022
-
[123]
A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by E nglish Marginal Abuse Models on T witter
Yee, Kyra and Schoenauer Sebag, Alice and Redfield, Olivia and Eck, Matthias and Sheng, Emily and Belli, Luca. A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by E nglish Marginal Abuse Models on T witter. Proceedings of the 3rd Workshop o...
2023
-
[124]
Examining the Causal Impact of First Names on Language Models: The Case of Social Commonsense Reasoning
Jeoung, Sullam and Diesner, Jana and Kilicoglu, Halil. Examining the Causal Impact of First Names on Language Models: The Case of Social Commonsense Reasoning. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023
2023
-
[125]
An Empirical Study of Metrics to Measure Representational Harms in Pre-Trained Language Models
Hosseini, Saghar and Palangi, Hamid and Awadallah, Ahmed Hassan. An Empirical Study of Metrics to Measure Representational Harms in Pre-Trained Language Models. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023
2023
-
[126]
Beyond T uring: A Comparative Analysis of Approaches for Detecting Machine-Generated Text
Adilazuarda, Muhammad. Beyond T uring: A Comparative Analysis of Approaches for Detecting Machine-Generated Text. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024
2024
-
[127]
Automated Adversarial Discovery for Safety Classifiers
Lal, Yash Kumar and Lahoti, Preethi and Sinha, Aradhana and Qin, Yao and Balashankar, Ananth. Automated Adversarial Discovery for Safety Classifiers. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024
2024
-
[128]
The Trade-off between Performance, Efficiency, and Fairness in Adapter Modules for Text Classification
Bui, Minh Duc and Von Der Wense, Katharina. The Trade-off between Performance, Efficiency, and Fairness in Adapter Modules for Text Classification. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024
2024
-
[129]
On the Interplay between Fairness and Explainability
Brandl, Stephanie and Bugliarello, Emanuele and Chalkidis, Ilias. On the Interplay between Fairness and Explainability. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (TrustNLP 2024). 2024
2024
-
[130]
F act A lign: Fact-Level Hallucination Detection and Classification Through Knowledge Graph Alignment
Rashad, Mohamed and Zahran, Ahmed and Amin, Abanoub and Abdelaal, Amr and Altantawy, Mohamed. F act A lign: Fact-Level Hallucination Detection and Classification Through Knowledge Graph Alignment. Proceedings of the 4th Workshop on Trustworthy Natural Language Processing (Trus...
2024
-
[131]
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refine
Kim, Heegyu and Cho, Hyunsouk. Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refine. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[132]
Ambiguity Detection and Uncertainty Calibration for Question Answering with Large Language Models
Shi, Zhengyan and Castellucci, Giuseppe and Filice, Simone and Kuzi, Saar and Kravi, Elad and Agichtein, Eugene and Rokhlenko, Oleg and Malmasi, Shervin. Ambiguity Detection and Uncertainty Calibration for Question Answering with Large Language Models. Proceedings of the 5th W...
2025
-
[133]
Error Detection for Multimodal Classification
Bonnier, Thomas. Error Detection for Multimodal Classification. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653/v1/2025.trustnlp-main.6
2025 doi
-
[134]
Know What You do Not Know: Verbalized Uncertainty Estimation Robustness on Corrupted Images in Vision-Language Models
Borszukovszki, Mirko and De Jong, Ivo Pascal and Valdenegro-Toro, Matias. Know What You do Not Know: Verbalized Uncertainty Estimation Robustness on Corrupted Images in Vision-Language Models. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[135]
Multi-lingual Multi-turn Automated Red Teaming for LLM s
Singhania, Abhishek and Dupuy, Christophe and Mangale, Shivam Sadashiv and Namboori, Amani. Multi-lingual Multi-turn Automated Red Teaming for LLM s. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[136]
Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
Kale, Sahil and Nadadur, Vijaykant. Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025
2025
-
[137]
M o N a C o: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Wolfson, Tomer and Trivedi, Harsh and Geva, Mor and Goldberg, Yoav and Roth, Dan and Khot, Tushar and Sabharwal, Ashish and Tsarfaty, Reut. M o N a C o: More Natural and Complex Questions for Reasoning Across Dozens of Documents. 2025. arXiv:2508.11133
2025 arXiv
-
[138]
and Aletras, Nikolaos and Ma, Ning
Hughes, Anthony and Duddu, Vasisht and Asokan, N. and Aletras, Nikolaos and Ma, Ning. PATCH : Mitigating PII Leakage in Language Models with Privacy-Aware Targeted Circuit Patc H ing. 2025. arXiv:2510.07452
2025
-
[139]
A Survey on Gender Bias in Natural Language Processing
Stanczak, Karolina and Augenstein, Isabelle. A Survey on Gender Bias in Natural Language Processing. 2021. arXiv:2112.14168
2021 arXiv
-
[140]
Inducing Positive Perspectives with Text Reframing
Ziems, Caleb and Li, Minzhi and Zhang, Anthony and Yang, Diyi. Inducing Positive Perspectives with Text Reframing. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. doi:10.18653/v1/2022.acl-long.257
2022 doi
-
[141]
The Importance of Modeling Social Factors of Language: Theory and Practice
Hovy, Dirk and Yang, Diyi. The Importance of Modeling Social Factors of Language: Theory and Practice. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. doi:10.18653/v1/2021.naa...
2021 doi
-
[142]
arXiv preprint arXiv:2307.09288 , year =
Llama 2: Open Foundation and Fine-Tuned Chat Models , author =. arXiv preprint arXiv:2307.09288 , year =
-
[143]
2023 , howpublished =
Model Card and Evaluations for Claude Models , author =. 2023 , howpublished =
2023
-
[144]
arXiv preprint arXiv:2303.08774 , year =
GPT-4 Technical Report , author =. arXiv preprint arXiv:2303.08774 , year =
-
[145]
Strength in Numbers: Estimating Confidence of Large Language Models by Prompt Agreement
Portillo Wightman, Gwenyth and Delucia, Alexandra and Dredze, Mark. Strength in Numbers: Estimating Confidence of Large Language Models by Prompt Agreement. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.t...
2023 doi
-
[146]
On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations
Cao, Yang Trista and Pruksachatkun, Yada and Chang, Kai-Wei and Gupta, Rahul and Kumar, Varun and Dhamala, Jwala and Galstyan, Aram. On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations. Proceedings of the 60th Annual Meeting o...
2022 doi
-
[147]
Pay Attention to the Robustness of C hinese Minority Language Models! Syllable-level Textual Adversarial Attack on T ibetan Script
Cao, Xi and Dawa, Dolma and Qun, Nuo and Nyima, Trashi. Pay Attention to the Robustness of C hinese Minority Language Models! Syllable-level Textual Adversarial Attack on T ibetan Script. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023...
2023 doi
-
[148]
Sample Attackability in Natural Language Adversarial Attacks
Raina, Vyas and Gales, Mark. Sample Attackability in Natural Language Adversarial Attacks. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.9
2023 doi
-
[149]
Detecting Personal Information in Training Corpora: an Analysis
Subramani, Nishant and Luccioni, Sasha and Dodge, Jesse and Mitchell, Margaret. Detecting Personal Information in Training Corpora: an Analysis. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.18
2023 doi
-
[150]
Privacy- and Utility-Preserving NLP with Anonymized data: A case study of Pseudonymization
Yermilov, Oleksandr and Raheja, Vipul and Chernodub, Artem. Privacy- and Utility-Preserving NLP with Anonymized data: A case study of Pseudonymization. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10.18653/v1/2023.trustnlp-1.20
2023 doi
-
[151]
Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models
Venkit, Pranav Narayanan and Srinath, Mukund and Wilson, Shomir. Automated Ableism: An Exploration of Explicit Disability Biases in Sentiment and Toxicity Analysis Models. Proceedings of the 3rd Workshop on Trustworthy Natural Language Processing (TrustNLP 2023). 2023. doi:10....
2023 doi
-
[152]
Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining
Cheng, Ruoxi and Ding, Yizhong and Cao, Shuirong and Wang, Zhiqiang and Shao, Shitong. Gibberish is All You Need for Membership Inference Detection in Contrastive Language-Audio Pretraining. Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025). 2025. doi:10.18653...
2025 doi
-
[153]
What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLM s
Abdelwahab, Mohamed and Collins, Michelle Yu and Chen, Sihan and Zhao, Yi Cheng and Mahmood, Zafarullah and Zhu, Jiading and Ali, Soliman and Rose, Jonathan. What are They Thinking? Delineation, Probing, and Tracking of Concepts in LLM s. Proceedings of the 6th Workshop on Tru...
2026 doi
-
[154]
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Al Ghussin, Yusser and Gurgurov, Daniil and Baeumel, Tanja and van Genabith, Josef and Schramowski, Patrick and Ostermann, Simon. Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection. Proceedings of the 6th Workshop on Trustworthy NL...
2026 doi
-
[155]
Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment
Bakman, Yavuz Faruk and Yaldiz, Duygu Nur and Avestimehr, Salman and Karimireddy, Sai Praneeth. Hair-Trigger Alignment: Black-Box Evaluation Cannot Guarantee Post-Update Alignment. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/202...
2026 doi
-
[156]
Domain-Dependent Safety Behavior in Open-Weight LLM s: An Empirical Study Across Seven Ethical Domains
Bugaud, Zacharie. Domain-Dependent Safety Behavior in Open-Weight LLM s: An Empirical Study Across Seven Ethical Domains. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.42
2026 doi
-
[157]
Single-Layer Activation Edits Easily Corrupt Factual Recall but Rarely Repair It
Bugaud, Zacharie. Single-Layer Activation Edits Easily Corrupt Factual Recall but Rarely Repair It. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.38
2026 doi
-
[158]
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
Mittal, Avni. Did You Forget What I Asked? Prospective Memory Failures in Large Language Models. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.33
2026 doi
-
[159]
Ghost Context: Measuring Cross-Context Interference in Long-Context Language Models
Namboothiri, Rohith. Ghost Context: Measuring Cross-Context Interference in Long-Context Language Models. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.19
2026 doi
-
[160]
The Geometry of Refusal: Linear Instability in Safety-Aligned LLM s
Ratnakar, Shivam and Vats, Kartikeya. The Geometry of Refusal: Linear Instability in Safety-Aligned LLM s. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.51
2026 doi
-
[161]
Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States
Sahoo, Subramanyam and Jain, Vinija and Chadha, Aman and Chaudhary, Divya. Linear Probes Detect Task Format, Not Reasoning Mode in Language Model Hidden States. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.trustnlp-main.12
2026 doi
-
[162]
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
Wang, Qianli and Wang, Mingyang and Feldhus, Nils and Ostermann, Simon and Cao, Yuan and Schuetze, Hinrich and M \"o ller, Sebastian and Schmitt, Vera. Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall. Proceedings of the 6th Works...
2026 doi
-
[163]
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
Xue, Zhiyu and Qi, Zimo and Liu, Guangliang and Chen, Bocheng and Pedarsani, Ramtin. Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment. Proceedings of the 6th Workshop on Trustworthy NLP ( T rust NLP 2026). 2026. doi:10.18653/v1/2026.t...
2026 doi
-
[164]
BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North A merican Chapter of the Association for Computational Linguistics: Hum...
2019 doi
-
[165]
GLUE : A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Wang, Alex and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R. GLUE : A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. Proceedings of the 2018 EMNLP Workshop B lackbox NLP : Analyzing and Interpreting Ne...
2018 doi
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.