Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:51:17.259117Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 118 outbound references and 0 inbound Pith citation observations for arXiv:2608.11171.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:51:17.259117Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 118 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8d9ca6f8-ec40-4cd2-9a0e-f359d768f7d3 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de088d98-e2f9-4e2d-af73-5700f1053308 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5ef2da-7092-4ecf-aada-16d42922fd86 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c011e2b4-367e-4497-9fa6-fa23c58f332f · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8e8bca1-0a47-4526-81e7-a4ff31fb9f57 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baac49ac-a54a-4be1-b053-273459137e61 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ffb726-c5f2-42d2-b97f-597a473e3724 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop don't forget the teachers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b16f8c3-3af4-417c-bddc-72aef843c6b5 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9870e611-b528-4ab9-9e31-792d137b9324 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop TrustLLM: Trustworthiness in Large Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46401178-93bd-4429-a47e-4e174c89416b · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831c3195-f441-40e5-a3e8-22ca521a3feb · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 163ffdfc-8028-4fbd-a4a9-71c18d0b6783 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a998e40-6912-4942-baef-7f9c23cbe17e · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 645cca4e-9f50-4fb8-93ce-98e5cd658204 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f74c9f77-3b20-4494-981d-04f072dc2f58 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e16acf-8f1e-42f5-a2a3-a5e70d8948fa · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6041a784-e8cb-42c8-aa8a-ad32cb84e29a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cddb846-25b5-4780-9965-db36f7c79c4b · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a32cee-7b52-4259-8f3a-8fe60d1019e6 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02bafdd9-53fb-4a9a-a152-1238249fb910 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c7f868c-804e-474f-8090-2f9f3b95bf1c · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8f1fc8-d6fe-4850-8274-79fed12bacdf · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db943417-ff45-4671-91b0-ed7794a7416a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3b4fb2d-528f-4e0a-89e8-0382adcebc0b · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c479cfcd-9596-4125-8055-ad17f421356b · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c9f417-3d2e-4107-bed1-25062866139e · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2026 , howpublished =
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f9553c-e3d4-432a-a788-02e54cf8dc1e · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Human-Centered Explainable AI : Towards a Reflective Sociotechnical Approach
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107975d5-4d78-4da4-8757-05817c127b83 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop and Wintersberger, Philipp and Manger, Carina and Hubig, Nina and Savage, Saiph and Weisz, Justin D
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e2ea7ae-a4f9-46c1-9066-b8b0cf35c35a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop A Survey on Medical Large Language Models: Technology, Application, Trustworthiness, and Future Directions
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da851883-7758-46cc-820b-29c1860d8b97 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Standard Benchmarks Fail -- Auditing LLM Agents in Finance Must Prioritize Risk
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91bc9c9d-60ba-4cbf-9cf8-667c86e04d82 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Don't Forget the Teachers
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 913fa84b-6be8-4080-8c9f-6d436179cc43 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2025 Silicon Valley Cybersecurity Conference (SVCC) , pages=
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bd9c0e2-d80d-419f-a421-d48a1d9b36af · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2023 , url =
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cf52372-36ab-4930-b5d2-9427461bb01c · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2024 , url=
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 428c3343-a31c-4b1a-a90f-2fc70ef11344 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efb6439-584c-4411-a97b-fb0b4b0a62ce · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2025 , publisher=
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 498ceeaf-2846-448f-bc20-7fc603cd67ae · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a8edc6-c58f-404c-b0a7-82c03f5dcfb3 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Interpretability Rules: Jointly Bootstrapping a Neural Relation Extractor with an Explanation Decoder
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa99828f-2299-4b74-8df4-932c0a7707fd · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Measuring Biases of Word Embeddings: What Similarity Measures and Descriptive Statistics to Use?
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0d500bcb-aa3f-47bf-870e-b15566842501 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop and Kiritchenko, Svetlana and Balkir, Esma
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 977ce2f1-f298-4342-b774-2ceae3df2885 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop GPT s Don ' t Keep Secrets: Searching for Backdoor Watermark Triggers in Autoregressive Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855102a8-6674-4972-b483-eab017c5bbfb · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Reliability Check: An Analysis of GPT -3's Response to Sensitive Topics and Prompt Wording
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2e7835e-f913-4a22-b938-6112a387e6b0 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Driving Context into Text-to-Text Privatization
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a569e7ad-2d73-44ca-8d2e-99b7c9e94b1e · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Expanding Scope: Adapting E nglish Adversarial Attacks to C hinese
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b3994e2-3d6e-410b-8a35-77b802cde3ba · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Flatness-Aware Gradient Descent for Safe Conversational AI
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a824ce0-ecf5-4508-b1c1-5e2749353908 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop PBI -Attack: Prior-Guided Bimodal Interactive Black-Box Jailbreak Attack for Toxicity Maximization
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70156afd-43f2-4b56-9937-21f1a44382fc · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Beyond Text-to- SQL for IoT Defense: A Comprehensive Framework for Querying and Classifying IoT Threats
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43c6f142-7da7-4d57-95c6-a3c650c01f21 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Minimal Evidence Group Identification for Claim Verification
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd32d46d-f950-4093-ad1d-ba9016f142d7 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Estimating Knowledge in Large Language Models Without Generating a Single Token
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f4c745-e219-4e72-80c3-ed429bbe5883 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Intrinsic Test of Unlearning Using Parametric Knowledge Traces
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fde013c0-e009-4fed-9000-b955d5bd1970 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Can we trust the evaluation on C hat GPT ?
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb79b921-8c81-4537-bcb9-a161c973fd60 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Improving Factuality of Abstractive Summarization via Contrastive Reward Learning
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f29426-d024-4297-bba4-f31f8c9d42e3 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Exploring Causal Mechanisms for Machine Text Detection Methods
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29cfc788-99b7-49f6-9547-553936eebf07 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop On the Robustness of Agentic Function Calling
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e27a5b00-a18a-48c3-8fc1-db57c08867f8 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Cross-Task Defense: Instruction-Tuning LLM s for Content Safety
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9e1eb20-7381-4116-8447-fdb5428c4c1e · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Gender Bias in Natural Language Processing Across Human Languages
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4df4eb7-a254-42fa-a742-7af291b2e8df · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Into the Gap between What Language Models Say and What They Know
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd145f3-52cc-4863-b58e-f57ea78259a9 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop The False Sense of Privacy in LLM s: Non-Verbatim Memorization and Semantic Leakage
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 665d41c0-4741-4d55-bfd6-a3f69d137367 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop and Raimundo, Marcos M
Reference 106
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df6bb622-cc75-4d30-a7c2-7f3f0cf08de8 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2023 , howpublished =
Reference 107
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8febde7-8661-42d9-a1ad-9d74840a86c0 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Nature Machine Intelligence , volume=
Reference 108
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f3659e-2cca-43b0-90de-d40d432325ad · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Formalizing Trust in Artificial Intelligence: Prerequisites, Causes and Goals of Human Trust in AI , year =
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ae1e4f8-d7f1-4291-852c-520c7d802d47 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop FAccT 2022 , year =
Reference 110
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0147a468-2552-4eda-a574-a7dee4c5c3ae · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop SODAPOP : Open-Ended Discovery of Social Biases in Social Commonsense Reasoning Models
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a4a719c5-83ac-4805-bb0e-d9e500235ea1 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop F air B elief - Assessing Harmful Beliefs in Language Models
Reference 112
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1b58e1-2ae7-4f1e-b79b-900131422633 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Investigating and Addressing Hallucinations of LLM s in Tasks Involving Negation
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c78b3f-8159-465f-af7a-ee3361fa9b89 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Introducing G en C eption for Multimodal LLM Benchmarking: You May Bypass Annotations
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae8427c-ff08-47d1-a8a4-5c26bd86b8ef · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Tell Me Why: Explainable Public Health Fact-Checking with Large Language Models
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006a2779-9e8f-4f0a-bcfe-076eb44a3cef · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Disentangling Linguistic Features with Dimension-Wise Analysis of Vector Embeddings
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb1c49fa-76ae-401e-b19e-adb654a3615f · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop On The Real-world Performance of Machine Translation: Exploring Social Media Post-authors' Perspectives
Reference 117
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e020f1c-d738-4c26-8411-4ed746e324a8 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop V i B e: A Text-to-Video Benchmark for Evaluating Hallucination in Large Multimodal Models
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525ff90f-5571-4d9d-8c82-0760ff712954 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop FACTOID : FAC tual en T ailment f O r halluc I nation Detection
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90472fb8-a17a-4039-83a5-b2ab7fd06287 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Private Release of Text Embedding Vectors
Reference 120
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74282c53-747a-475f-98ee-858b036da7a9 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Challenges in Applying Explainability Methods to Improve the Fairness of NLP Models
Reference 121
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb8e51e-980a-4099-8f0a-675e9c86307c · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop An Encoder Attribution Analysis for Dense Passage Retriever in Open-Domain Question Answering
Reference 122
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b90ff77c-5bbe-4c31-85ed-7a757a7e202a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop A Keyword Based Approach to Understanding the Overpenalization of Marginalized Groups by E nglish Marginal Abuse Models on T witter
Reference 123
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adf32ab9-b7de-43be-b04c-19f465a8f73d · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Examining the Causal Impact of First Names on Language Models: The Case of Social Commonsense Reasoning
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9fa36f13-191e-494a-b98d-78ff65c182fa · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop An Empirical Study of Metrics to Measure Representational Harms in Pre-Trained Language Models
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 341e203f-54ec-4c36-8eb9-9c88008c453a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Beyond T uring: A Comparative Analysis of Approaches for Detecting Machine-Generated Text
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 005475b9-56ca-431d-848e-697fccbee092 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Automated Adversarial Discovery for Safety Classifiers
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7ea20418-09e4-482a-9562-f91ede53ab8f · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop The Trade-off between Performance, Efficiency, and Fairness in Adapter Modules for Text Classification
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 190fa7b7-a3c4-4d80-876e-5395f7ffa0f7 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop On the Interplay between Fairness and Explainability
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55ba8d40-09b8-46fd-96da-4a6b551689e9 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop F act A lign: Fact-Level Hallucination Detection and Classification Through Knowledge Graph Alignment
Reference 130
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c75aa3a-5794-4c63-88b4-404d2bc63ee8 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refine
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f66e172-2292-46bb-b6a1-953db8cc11f4 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Ambiguity Detection and Uncertainty Calibration for Question Answering with Large Language Models
Reference 132
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 29146040-7016-4cff-8c6c-ea063beffc48 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Error Detection for Multimodal Classification
Reference 133
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91c8a12d-937d-4e08-ba69-3eec055fa2a8 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Know What You do Not Know: Verbalized Uncertainty Estimation Robustness on Corrupted Images in Vision-Language Models
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a7212c6c-36cd-4ffb-b00b-7e7f9fc293b7 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Multi-lingual Multi-turn Automated Red Teaming for LLM s
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5be57346-eaa7-4f20-8844-ef8f0db1a5bf · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries
Reference 136
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4fb8417f-d540-425e-9264-e7743f4588b2 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents
Reference 137
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f6a97a-3cd1-4e93-afd0-4e4e809588cf · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop and Aletras, Nikolaos and Ma, Ning
Reference 138
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7046bf1a-afc2-43c3-b9cf-0cafdae14f5b · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop A Survey on Gender Bias in Natural Language Processing
Reference 139
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 443dba07-8f5f-40a6-8d92-ae7ad2edc7fa · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Inducing Positive Perspectives with Text Reframing
Reference 140
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a71ace5c-74d1-4d18-9a9d-2fe24ea8f402 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop The Importance of Modeling Social Factors of Language: Theory and Practice
Reference 141
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e9ce44b-5ca1-4dc6-b711-3998d2ef9c6d · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 142
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 258c868f-edfa-44fd-8a50-a06277381181 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop 2023 , howpublished =
Reference 143
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7650e705-1670-4fde-aa0c-ef1a3c49304a · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop GPT-4 Technical Report
Reference 144
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a4c4ba2-69bc-47c8-92bb-119dbf2d7bf4 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Strength in Numbers: Estimating Confidence of Large Language Models by Prompt Agreement
Reference 145
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4617b8f3-ce41-4840-a759-dd30de6748b9 · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop On the Intrinsic and Extrinsic Fairness Evaluation Metrics for Contextualized Language Representations
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b6dd65d-2dd3-4dc9-b107-a6e7c1e3b72c · outbound
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop Pay Attention to the Robustness of C hinese Minority Language Models! Syllable-level Textual Adversarial Attack on T ibetan Script
Reference 147
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.