REVIEW 4 major objections 4 minor 48 references
RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?
T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read RV4Chatbot monitors intent-based chatbots at runtime, verifying each intent and action against formal interaction protocols.
desk verdict Solid, reproducible engineering: a general RV framework for intent-based chatbots with two instantiations, but the safety claim rests on unverified NLU accuracy and a single 12-message benchmark. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the decision wrapper, a thin instrumentation layer inside the chatbot's decision maker. It intercepts two kinds of events—user intents with their parameters (after NLU classification) and chatbot actions before execution—and forwards them to an external runtime monitor. The monitor, which only needs to output true, false, or inconclusive verdicts, checks the event stream against an interaction protocol formalised in a specification language such as RML, a domain-specific language for parametric, potentially non-context-free properties. A false verdict returns the chatbot to a listening state with an error message; true or inconclusive verdicts leave the flow unchanged. This wrapper is the single point of instrumentation, which is what makes the framework formalism-agnostic and minimally invasive across different chatbot frameworks.
What would settle it
Run a conversation in which the NLU is deliberately misled—for example, a request to add an object at an already occupied position phrased so that the intent classifier labels it as a removal request—and check whether the monitor catches the resulting unsafe action. If the violation passes unnoticed, the framework's safety guarantee is contingent on NLU accuracy rather than on the protocol itself. A different falsifier would be to measure response time under hundreds of concurrent monitored conversations; if overhead grows superlinearly or dominates conversation latency, the negligible-overhead claim does not scale.
Extended reading notes
Core claim
On the paper's own terms, the central claim is that the correctness of a user–chatbot conversation can be assessed as a coherent whole by monitoring the stream of user intents and chatbot actions, rather than by inspecting how individual messages are generated. To that end, RV4Chatbot provides a logical architecture—a decision wrapper inside the chatbot that forwards events to an external monitor—that is parametric in both the chatbot framework and the monitor's specification language. Two concrete instantiations demonstrate the approach: RV4Rasa, realised by adding a monitor policy to Rasa's policy stack, and RV4Dialogflow, realised by an instrumentation script that generates a policy component forwarding messages to a webhook monitor. In both cases the monitor emits boolean verdicts; a false verdict routes the conversation to an error state, so no unsafe action is executed. The paper also shows that the same instrumented chatbot can be used offline for testing and then at runtime, with no code changes, and reports experiments in which monitoring overhead on a twelve-message conversation is negligible.
Load-bearing premise
The monitor's verdicts are only as trustworthy as the chatbot's intent classifier and entity extractor, because the monitor sees the intents and slots produced by the NLU component rather than the user's raw words; if the NLU misclassifies a request, the monitor will check a property against the wrong event.
Editorial extensions
If this is right
- The same decision-wrapper design can be ported to any intent-based chatbot framework that exposes its intents and actions to a policy or webhook layer, not just Rasa and Dialogflow.
- A chatbot instrumented for runtime verification can be reused for offline testing by swapping the human user for a scripted sentence generator, with no changes to the monitor or the instrumented code.
- Safety properties that depend on conversation history, such as 'do not add an object to an already occupied position,' can be enforced at runtime, preventing unsafe actions before they are executed.
- Because the framework is formalism-agnostic, properties can be written in any runtime-verification language that can emit true/false/inconclusive verdicts, including languages more expressive than LTL.
Reading between the lines
- The framework implicitly shifts the safety burden to NLU quality: it verifies protocol compliance of the interpreted conversation, not of the user's actual words, so adversarial or noisy inputs that fool intent classification would bypass the monitor. A testable extension would be to feed the raw user text to the monitor as an additional event and verify consistency between the NLU output and the
- The same architecture could be adapted to generative chatbots by defining coarse-grained observable events—such as calls to external tools or explicit safety-relevant decisions—and verifying protocols over those events, though the authors leave this for future work.
- The negligible-overhead measurement covers a single 12-message conversation; the modularity intuition suggests scalability, but the paper does not test many concurrent conversations, so a stress test with realistic message volumes would be the natural next experiment.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RV4Chatbot, a runtime verification framework for intent-based chatbots. Expected chatbot behavior is formalized as interaction protocols between the user and the chatbot, and a monitor checks the stream of recognized intents, extracted parameters, and chatbot actions against properties written in the Runtime Monitoring Language (RML). Two instantiations are presented: RV4Rasa, which adds a monitor policy to Rasa, and RV4Dialogflow, which instruments a Dialogflow agent through a generated policy component. The paper reports a factory-automation case study with three safety properties and a performance experiment claiming that the monitor introduces negligible overhead.
Significance. If the claims hold, the paper makes a useful contribution: it provides a concrete, open-source architecture for runtime monitoring of conversational AI chatbots, with two independent instantiations and a clear separation between the monitor, the decision wrapper, and the chatbot framework. The use of RML allows parametric, protocol-level properties that go beyond simple intent checks, and the choice to make the framework formalism-agnostic is a genuine strength. The open-source artifacts (RV4Rasa and RV4Dialogflow code) are a positive aspect that supports reproducibility. However, the significance is currently limited by the thinness of the experimental evaluation and by an unexamined assumption about the reliability of the NLU component, both of which are load-bearing for the paper's safety and overhead claims.
major comments (4)
- [Section 7.3, Figure 6] The claim that monitoring introduces negligible overhead is not supported by the reported data. Figure 6 shows times for a single 12-message test conversation, with no indication of the number of repetitions, no variance or confidence intervals, and no hardware/software environment details. Although Section 7.3 states that run_test.py iterates the conversation 'a certain number of times', the number and the per-iteration statistics are never reported. Without this information, the apparent lack of overhead could be due to measurement noise or to the particular short conversation chosen. Please provide the number of runs, the mean and dispersion per message, and a description of the experimental environment, and compare the real-monitor condition against both the no-monitor and dummy-monitor baselines.
- [Section 4, Figure 1; Section 7.2] The monitor observes the recognized intents and slots produced by the NLU component (event 3 in Figure 1), not the user's raw utterance, and the RML properties in Section 7.2 are defined over those NLU outputs, e.g., add_object(x,y) matches {intent:{name:'add_object'}, slots:{horizontal:x, vertical:y}}. Consequently, if the NLU misclassifies the user's request or extracts incorrect parameters, the monitor will evaluate the property on a trace that does not correspond to the user's actual request, and an unsafe request can pass. The paper explicitly disclaims responsibility for message-level correctness in Section 2 ('our focus is not on whether the model correctly produces or classifies individual messages'), and the only NLU-related property, the confidence > 60% check in Section 7.2, does not entail correctness. This is a load-bearing premise for the safety claims; please state the guarantee precisely as conditional on NLU accuracy and, ideally, evaluate the framework under NLU misclassification.
- [Abstract and Section 4 vs. Section 5.2] The paper's safety claims are stronger than what the Rasa instantiation delivers. The abstract promises that chatbots 'consistently adhere to expected, safe behaviours', and Section 4 states that after a false verdict 'no unsafe actions are performed', implying prevention. Section 5.2, however, says that in RV4Rasa the monitorPolicy 'can only stop the chatbot immediately after the wrong action has been executed' and that the authors 'give up prevention', accepting ex-post notification. This is a substantive discrepancy between the general architecture and one of the two reported instantiations. The paper should either qualify the general claims as reactive rather than preventive, or explain which instantiations achieve prevention and which do not.
- [Section 7, introductory paragraph] The statement that 'the monitor always works as expected' and the claim that all properties 'are correctly verified by the monitor' are assertions without supporting evidence. No qualitative evaluation is reported: there are no example traces showing violations being detected, no counts of true positives, false positives, or false negatives, and no comparison of monitor verdicts against an independent oracle. Since the correctness of the monitor is central to the framework's value, this assertion should be backed by a concrete evaluation protocol and results, or the claims should be scaled back accordingly.
minor comments (4)
- [Section 7.2] In the RML specification for the AddObject property, the term uses 'not_add_ob ject(x,y)' (with a typographical space and no definition), but the event types list ETs does not include this event type. Please add the definition or correct the typo.
- [Section 5.3] The code block showing the policy configuration has inconsistent spacing, a line break inside 'policies :', and the policy class appears both as 'monitorPolicy' and 'MonitorPolicy'; please harmonize the presentation.
- [Figure 6] Figure 6 lacks axis labels and a legend explaining what the plotted values represent; please add units for the time axis and a description of the bars.
- [Throughout] The paper uses 'DialogFlow' and 'Dialogflow' interchangeably; please choose one spelling and apply it consistently.
Circularity Check
No circular derivation: the framework's instantiations and overhead measurements are independent of the properties they check.
full rationale
The derivation chain is self-contained. RV4Chatbot's central claims—architecture, two instantiations, and negligible overhead—are supported by concrete implementations (monitor_policy.py, instrumenter.py), public code, and measurements comparing no monitor, dummy monitor, and real monitor. RML is cited from the authors' prior work, but it is used as one instantiation language and the framework is explicitly parametric in the RV language; no property, parameter, or result is fitted to predetermine the outcome. The safety properties are formal requirements from the factory scenario (ISO 10218), not extracted from the monitor's verdicts. The paper's disclaimer that it does not verify NLU classification (Section 2) and the absence of qualitative violation experiments (Section 7) are limitations and missing evidence, not circularity: the monitor checks the recognized intent/action trace, which is a modeling boundary acknowledged in the text. Self-citations to [22], [24], and the RML papers are routine and not load-bearing.
Assumptions & free parameters
free parameters (1)
- confidence_threshold =
60%
assumptions (3)
- domain assumption The event stream sent by the decision wrapper (recognized intents and bot actions) faithfully represents the actual user request and the actual effect of the chatbot's action on the environment.
- standard math A runtime monitor can be defined that reads one event at a time and emits a verdict in {true, false, inconclusive}, and this is sufficient for the RV4Chatbot recovery logic.
- domain assumption RML is expressive enough to encode the safety properties and its semantics, as published in [6], are correct.
Cite this review
Pith. "Pith review of RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?." pith.science (2026). https://pith.science/paper/P6YK4VVN
@misc{pith2026241114368,
author = {Pith},
title = {Pith review of: RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?},
year = {2026},
howpublished = {\url{https://pith.science/paper/P6YK4VVN}},
note = {Machine review of arXiv:2411.14368}
}
read the original abstract
Chatbots have become integral to various application domains, including those with safety-critical considerations. As a result, there is a pressing need for methods that ensure chatbots consistently adhere to expected, safe behaviours. In this paper, we introduce RV4Chatbot, a Runtime Verification framework designed to monitor deviations in chatbot behaviour. We formalise expected behaviours as interaction protocols between the user and the chatbot. We present the RV4Chatbot design and describe two implementations that instantiate it: RV4Rasa, for monitoring chatbots created with the Rasa framework, and RV4Dialogflow, for monitoring Dialogflow chatbots. Additionally, we detail experiments conducted in a factory automation scenario using both RV4Rasa and RV4Dialogflow.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Abubakar Abid, Maheen Farooqi & James Zou (2021): Persistent Anti-Muslim Bias in Large Language Models. In: AIES, ACM, pp. 298–306, doi:10.1145/3461702.3462624
arXiv 2021
-
[2]
Ma- chine Learning with Applications 2, p
Eleni Adamopoulou & Lefteris Moussiades (2020): Chatbots: History, technology, and applications . Ma- chine Learning with Applications 2, p. 100006, doi:10.1016/j.mlwa.2020.100006
arXiv 2020
-
[3]
Hind Alotaibi & Hussein Zedan (2010): Runtime verification of safety properties in multi-agents systems. In: 10th International Conference on Intelligent Systems Design and Applications, ISDA 2010, November 29 - December 1, 2010, Cairo, Egypt, IEEE, pp. 356–362, doi:10.1109/ISDA.2010.5687238
-
[4]
Available at https://rmlatdibris.github.io/
Davide Ancona, Angelo Ferrando, Luca Franceschini & Viviana Mascardi: RML web site . Available at https://rmlatdibris.github.io/. Accessed on November 22, 2024
work page 2024
-
[5]
In: Theory and Practice of Formal Methods, LNCS 9660, Springer, pp
Davide Ancona, Angelo Ferrando & Viviana Mascardi (2016): Comparing Trace Expressions and Linear Temporal Logic for Runtime Verification. In: Theory and Practice of Formal Methods, LNCS 9660, Springer, pp. 47–64, doi:10.1007/978-3-319-30734-3_6
-
[6]
Davide Ancona, Luca Franceschini, Angelo Ferrando & Viviana Mascardi (2021): RML: Theory and practice of a domain specific language for runtime verification . Sci. Comput. Program. 205, p. 102610, doi:10.1016/j.scico.2021.102610. 88 RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?
arXiv 2021
-
[7]
Najwa Abu Bakar & Ali Selamat (2013): Runtime Verification of Multi-agent Systems Interaction Quality . In: Intelligent Information and Database Systems - 5th Asian Conf., ACIIDS 2013 , LNCS 7802, Springer, Berlin, Heidelberg, pp. 435–444, doi:10.1007/978-3-642-36546-1_45
-
[8]
In: Lectures on Runtime Verification - Introductory and Advanced Topics, LNCS 10457, Springer, pp
Ezio Bartocci, Yliès Falcone, Adrian Francalanza & Giles Reger (2018):Introduction to Runtime Verification. In: Lectures on Runtime Verification - Introductory and Advanced Topics, LNCS 10457, Springer, pp. 1–33, doi:10.1007/978-3-319-75632-5_1
Show all 48 references
-
[9]
computers & security 29(3), pp
Andreas Bauer & Jan Jürjens (2010): Runtime verification of cryptographic protocols. computers & security 29(3), pp. 315–330, doi:10.1016/j.cose.2009.09.003
2010 doi
-
[10]
In: DIMACS/SYCON WS on Verification and Control of Hybrid Systems, LNCS 1066, Springer, pp
Johan Bengtsson, Kim Guldstrand Larsen, Fredrik Larsson, Paul Pettersson & Wang Yi (1995): UPPAAL - a Tool Suite for Automatic Verification of Real-Time Systems. In: DIMACS/SYCON WS on Verification and Control of Hybrid Systems, LNCS 1066, Springer, pp. 232–243, doi:10.1007/BFB0020949
1995 doi
- [11]
-
[12]
Available at https://botium-docs.readthedocs.io/en/latest/
Botium: Bots Testing Bots . Available at https://botium-docs.readthedocs.io/en/latest/. Ac- cessed on November 22, 2024
2024
-
[13]
Josip Bozic (2022): Ontology-based metamorphic testing for chatbots . Softw. Qual. J. 30(1), pp. 227–251, doi:10.1007/s11219-020-09544-9
2022 doi
-
[14]
Tazl & Franz Wotawa (2019): Chatbot Testing Using AI Planning
Josip Bozic, Oliver A. Tazl & Franz Wotawa (2019): Chatbot Testing Using AI Planning. In: IEEE Int. Conf. On Artificial Intelligence Testing, AITest 2019, IEEE, pp. 37–44, doi:10.1109/AITest.2019.00-10
2019 doi
-
[15]
In: Testing Software and Systems - 31st IFIP WG 6.1 Int
Josip Bozic & Franz Wotawa (2019): Testing Chatbots Using Metamorphic Relations. In: Testing Software and Systems - 31st IFIP WG 6.1 Int. Conf., ICTSS 2019, LNCS 11812, Springer, pp. 41–55, doi:10.1007/978- 3-030-31280-0_3
2019 doi
-
[16]
In: Quality of Information and Communications Technology - 13th Int
Sergio Bravo-Santos, Esther Guerra & Juan de Lara (2020): Testing Chatbots with Charm. In: Quality of Information and Communications Technology - 13th Int. Conf., QUATIC 2020 , CCIS 1266, Springer, pp. 426–438, doi:10.1007/978-3-030-58793-2_34
2020 doi
-
[17]
CoRR abs/2310.15469, doi:10.48550/ARXIV .2310.15469
Xiaoyi Chen, Siyuan Tang, Rui Zhu, Shijun Yan, Lei Jin, Zihao Wang, Liya Su, XiaoFeng Wang & Haixu Tang (2023): The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks. CoRR abs/2310.15469, doi:10.48550/ARXIV .2310.15469. arXiv:2310.15469
-
[18]
Available at https://www.ibm.com/blog/chatbot-types/
Bella Church (2023): 5 types of chatbot and how to choose the right one for your business . Available at https://www.ibm.com/blog/chatbot-types/. Accessed on November 22, 2024
2023
-
[19]
Edmund M Clarke (1997): Model checking. In: Int. Conf. on Foundations of Software Technology and Theoretical Computer Science, Springer, pp. 54–56, doi:10.1007/BFb0058022
1997 doi
-
[20]
Robotics 12(2), p
Debora C Engelmann, Angelo Ferrando, Alison R Panisson, Davide Ancona, Rafael H Bordini & Viviana Mascardi (2023): RV4JaCa — Towards Runtime Verification of Multi-Agent Systems and Robotic Applica- tions. Robotics 12(2), p. 49, doi:10.3390/robotics12020049
2023 doi
-
[21]
Available at https://www.europarl.europa.eu/news/en/press-room/20231206IPR15699/ artificial-intelligence-act-deal-on-comprehensive-rules-for-trustworthy-ai
European Parliament (2023): Artificial Intelligence Act . Available at https://www.europarl.europa.eu/news/en/press-room/20231206IPR15699/ artificial-intelligence-act-deal-on-comprehensive-rules-for-trustworthy-ai . Ac- cessed on November 22, 2024
2023
-
[22]
In: 6th Int
Angelo Ferrando, Andrea Gatti & Viviana Mascardi (2023): RV4Rasa: A Formalism-Agnostic Runtime Ver- ification Framework for Verifying ChatBots in Rasa . In: 6th Int. WS on Verification and Monitoring at Runtime Execution, VORTEX 2023, ACM, pp. 1–8, doi:10.1145/3605159.3605855
2023
-
[23]
In Svetlana S
Asbjørn Følstad, Marita Skjuve & Petter Bae Brandtzæg (2018): Different Chatbots for Different Purposes: Towards a Typology of Chatbots to Understand Interaction Design . In Svetlana S. Bodrunova, Olessia Koltsova, Asbjørn Følstad, Harry Halpin, Polina Kolozaridi, Leonid Yulda...
2018
-
[24]
Robotics 12(2), p
Andrea Gatti & Viviana Mascardi (2023): VEsNA, a Framework for Virtual Environments via Natural Language Agents and Its Application to Factory Automation . Robotics 12(2), p. 46, doi:10.3390/ROBOTICS12020046
2023 doi
-
[25]
Geovana Ramos, Nunes R
Sousa S. Geovana Ramos, Nunes R. Genaína & Dias C. Edna (2023):A Modeling Strategy for the Verification of Context-Oriented Chatbot Conversational Flows via Model Checking . Journal of Universal Computer Science 29(7), pp. 805–835, doi:10.3897/jucs.91311
2023 doi
-
[26]
– GII (2024): Global Large Language Model (LLM) Market Research Report
Global Information, Inc. – GII (2024): Global Large Language Model (LLM) Market Research Report . Available at https://www.giiresearch.com/report/ qyr1384359-global-large-language-model-llm-market-research.html . Accessed on November 22, 2024
2024
-
[27]
Google: DialogFlow: Online Resource, https: // cloud. google. com/ dialogflow/. Available at https://cloud.google.com/dialogflow/
-
[28]
Available at https://cloud.google.com/dialogflow
Google: Dialogflow web site . Available at https://cloud.google.com/dialogflow. Accessed on November 22, 2024
2024
-
[29]
Available at https://gemini.google.com/
Google: Gemini web site. Available at https://gemini.google.com/. Accessed on November 22, 2024
2024
-
[30]
Available at https://cobusgreyling.medium
Cobus Greyling (2023): Conversational UIs & LLMs . Available at https://cobusgreyling.medium. com/large-language-model-llm-disruption-of-chatbots-8115fffadc22 . Accessed on Novem- ber 22, 2024
2023
-
[31]
Available at https://www.jasper.ai/chat
Jasper AI: Jasper web site. Available at https://www.jasper.ai/chat. Accessed on November 22, 2024
2024
-
[32]
Jaeho Jeon, Seongyong Lee & Hohsung Choe (2023): Beyond ChatGPT: A conceptual framework and systematic review of speech-recognition chatbots for language learning . Comput. Educ. 206, p. 104898, doi:10.1016/j.compedu.2023.104898
2023
-
[33]
Sun (2023): Gender bias and stereotypes in Large Language Mod- els
Hadas Kotek, Rikker Dockum & David Q. Sun (2023): Gender bias and stereotypes in Large Language Mod- els. In: The ACM Collective Intelligence Conf., CI 2023, ACM, pp. 12–24, doi:10.1145/3582269.3615599
2023
-
[34]
In: Australian Software Engineering Conf
Zheng Li, Yan Jin & Jun Han (2006): A runtime monitoring and validation framework for web service interactions . In: Australian Software Engineering Conf. (ASWEC’06) , IEEE, pp. 10–pp, doi:10.1109/ASWEC.2006.6
2006 doi
-
[35]
disaffordances
Xiaolin Lin, Bin Shao & Xuequn Wang (2022): Employees’ perceptions of chatbots in B2B marketing: Affordances vs. disaffordances . Industrial Marketing Management 101, pp. 45–56, doi:10.1016/j.indmarman.2021.11.016. Available at https://www.sciencedirect.com/science/ article/pi...
2022 doi
-
[36]
Loveland (1978): Automated theorem proving: a logical basis
Donald W. Loveland (1978): Automated theorem proving: a logical basis. Fundamental studies in computer science 6, North-Holland
1978
-
[37]
Available at https://www.marketsandmarkets
MarketsandMarkets (2023): Conversational AI Market. Available at https://www.marketsandmarkets. com/Market-Reports/conversational-ai-market-49043506.html . Accessed on November 22, 2024
2023
-
[38]
Available at https://wit.ai/
Meta: Wit.ai web site. Available at https://wit.ai/. Accessed on November 22, 2024
2024
-
[39]
Martin Mitrevski (2018): Getting started with wit.ai, doi:10.1007/978-1-4842-3396-2_5
2018 doi
-
[40]
Available at https://openai.com/blog/chatgpt
Open AI (2022): Introducing ChatGPT. Available at https://openai.com/blog/chatgpt. Accessed on November 22, 2024
2022
-
[41]
Available at https://rasa.com/
Rasa technologies: Rasa web site. Available at https://rasa.com/. Accessed on November 22, 2024
2024
-
[42]
90 RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?
Navin Sabharwal, Amit Agrawal, Navin Sabharwal & Amit Agrawal (2020): Introduction to Google Di- alogflow. 90 RV4Chatbot: Are Chatbots Allowed to Dream of Electric Sheep?
2020
-
[43]
Seshia, Ankush Desai, Tommaso Dreossi, Daniel J
Sanjit A. Seshia, Ankush Desai, Tommaso Dreossi, Daniel J. Fremont, Shromona Ghosh, Edward Kim, Sumukh Shivakumar, Marcell Vazquez-Chanlatte & Xiangyu Yue (2018): Formal Specification for Deep Neural Networks. In: Automated Technology for Verification and Analysis - 16th Int. ...
2018 doi
-
[44]
In: IEEE International Conference on Cloud Computing, CLOUD 2010, Miami, FL, USA, 5-10 July, 2010, IEEE Computer Society, pp
Jin Shao, Hao Wei, Qianxiang Wang & Hong Mei (2010): A Runtime Model Based Monitoring Approach for Cloud. In: IEEE International Conference on Cloud Computing, CLOUD 2010, Miami, FL, USA, 5-10 July, 2010, IEEE Computer Society, pp. 313–320, doi:10.1109/CLOUD.2010.31
2010 doi
-
[45]
Standard
Technical Committee:ISO/TC 299 Robotics (2011): Robots and robotic devices – Safety requirements for industrial robots. Standard
2011
-
[46]
Petzold, William Yang Wang, Xun Zhao & Dahua Lin (2023): Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda R. Petzold, William Yang Wang, Xun Zhao & Dahua Lin (2023): Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models . CoRR abs/2310.02949, doi:10.48550/ARXIV .2310.02949. arXiv:2310.02949
- [47]
-
[48]
(2023): Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Yue Zhang, Yafu Li, Leyang Cui & et al. (2023): Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv:2309.01219
2023 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.