REVIEW 5 major objections 5 minor 1 cited by
SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector
T0 review · 5 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Four-step framework turns public-sector AI risks into test prompts
desk verdict A clear, well-scoped proposal for a public-sector AI risk evaluation pipeline whose central claims about systematicity and comprehensiveness are not yet backed by any implementation or data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is SAIF, a four-stage prompt-data generation pipeline. It combines a risk taxonomy with scenario templates, a fixed set of jailbreak methods, and a catalog of prompt types to produce evaluation prompts; the work it does is converting broad risk categories into concrete, machine-testable inputs that can be reused across models and modalities.
What would settle it
Apply SAIF to a known public-sector risk subtopic (for example, hateful content in a welfare chatbot) using only its listed jailbreak methods, and compare with a broader attack library that includes encoding-based or multi-turn jailbreaks; if the broader library reliably finds refusals or unsafe outputs that SAIF's methods miss, then the framework's claim to enable comprehensive evaluation fails for that subtopic.
Extended reading notes
Core claim
SAIF claims to make generative-AI risk evaluation in the public sector systematic by decomposing it into four stages: breaking down risks, designing scenarios, applying jailbreak methods, and exploring prompt types. The first stage selects subtopics within four risk factors (system and operational misuse, content safety, societal, and legal and rights-related), drawing on a taxonomy built from government policies and corporate guidelines. The second designs modality-specific scenarios so that text, image, and video generations are all covered. The third applies refusal suppression, disguised intent, and hypothetical-scenario jailbreaks to test whether safeguards hold under attack. The fourth varies the expression of the attack through prompt types such as chain-of-thought, role-playing, expert prompting, rails, and reflection. The resulting prompt data is meant to be fed to LLMs and LMMs and scored on a Likert scale for the targeted risks, giving a quantifiable, comparable vulnerability profile.
Load-bearing premise
The framework's 'comprehensive' coverage rests on the manually chosen subtopics in Table 1 and on the three jailbreak methods and handful of prompt types selected—if those miss a risk or an attack style, the evaluation will show no vulnerability even though one exists.
Editorial extensions
If this is right
- The four-stage pipeline can be reused each time a new risk factor or modality appears, so evaluations do not have to be rebuilt from scratch.
- Because prompt data is produced consistently, results across models and across update versions become comparable on the same risk dimensions.
- The multimodal scenarios let public-sector risk evaluators probe image and video generation, not just text, which matches how governments actually use these tools.
- The framework's modular design means new jailbreak methods and prompt types can be added without changing the risk taxonomy.
Reading between the lines
- The paper does not report coverage analysis for its subtopics; a natural next step would be to check whether Table 1's subtopics saturate the risk taxonomy or are biased toward conspicuous risks like hateful content over quieter ones like subtle discrimination.
- SAIF's completeness claim could be made empirically testable by measuring whether adding more jailbreak methods or prompt types reduces the number of risk-tagged outputs—if it doesn't, the selected set is likely sufficient.
- The same four-stage structure could be transferred outside the public sector to regulated domains such as healthcare or finance, where the risk taxonomy would need to be swapped but the scenario-design and jailbreak stages carry over.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SAIF, a four-stage framework for generating prompt data to evaluate generative AI risks in the public sector. It builds on an existing risk taxonomy (Zeng et al. 2024), revisits that taxonomy in the public-sector context, and extends it to text, image, and video modalities. The four stages are: breaking down risks into subtopics, designing scenarios, applying jailbreak methods, and exploring prompt types. The paper claims that this pipeline ensures systematic and consistent generation of prompt data and facilitates comprehensive evaluation, providing a foundation for mitigation. The manuscript contains no experiments, no generated dataset, no model outputs, and no formal or empirical validation; it is primarily a framework proposal with illustrative examples in Figures 1 and 2 and a subtopic list in Table 1.
Significance. If validated, SAIF could offer a useful structured approach to generating risk-evaluation prompts for public-sector generative AI, and the multimodal extension is timely given the increasing deployment of image and video generators. The paper's main strengths are its explicit decomposition of the prompt-generation process into four stages and its reuse of an externally derived, policy-grounded risk taxonomy, which gives the framework a credible starting point. I do not see a circularity problem: the framework has no fitted parameters, and the taxonomy is external to the paper. However, the central claims of systematicity, consistency, and comprehensiveness are entirely unverified. No coverage analysis, selection criteria, comparison to alternative frameworks, or empirical demonstration is provided, and the future-work section concedes that the current subtopic selection is not yet diverse or rigorous. The contribution is therefore a plausible organizational scheme rather than a demonstrated evaluation framework; the paper's current title and abstract overstate what is shown.
major comments (5)
- [Abstract; Sec. 'Systematic Data Generation Framework'] The claim that SAIF 'ensures the systematic and consistent generation of prompt data, facilitating a comprehensive evaluation' is the central assertion of the paper, but it is not supported by any experimental evidence, formal definition, or coverage analysis. The paper presents no generated dataset, no model outputs, no human-annotation results, and no comparison to alternative risk taxonomies or prompt-generation methods. The four stages are described at a high level with illustrative examples, leaving the reader unable to verify whether the pipeline actually produces systematic, consistent, or comprehensive prompt data. This is a load-bearing unsupported claim that should be either demonstrated or substantially weakened.
- [Sec. 'Breaking Down Risks'; Table 1; Conclusion and Future Work] The comprehensiveness of the evaluation depends on the completeness and representativeness of the manually curated subtopics in Table 1, but the paper gives no selection criteria, coverage metric, or justification for why these subtopics span the space of public-sector risks. The future-work section states that knowledge graphs will enable 'a more diverse and rigorous exploration of risk-related subtopics,' which implicitly concedes that the current subtopic set is not yet sufficiently diverse or rigorous. Since the 'comprehensive evaluation' claim rests directly on this list, the paper should either provide a principled method for subtopic selection and a demonstration of coverage, or explicitly limit the claim to the listed subtopics.
- [Sec. 'Designing Scenarios'; Figure 1] The multimodal extension is asserted rather than demonstrated. The 'Designing Scenarios' section offers only illustrative examples ('a scenario for a text-based modality might involve...') with no procedure for translating a subtopic into a scenario, and no operational definition of when a set of scenarios is comprehensive for a given modality. For image and video modalities, the only concrete artifacts shown in Figure 1 are textual prompts; no image- or video-generation and evaluation pipeline is specified or validated. As written, the paper does not show that the framework applies to modalities beyond text in any reproducible way.
- [Secs. 'Applying Jailbreak Methods' and 'Exploring Prompt Types'] The selection of three jailbreak methods (refusal suppression, disguised intent, hypothetical scenario) and six prompt types (CoT, zero-shot CoT, role-playing, expert prompting, rails, reflection) is presented without any criteria for comprehensiveness or representativeness. No evidence is provided that this particular combination spans the attack surface for generative AI in the public sector, nor that these methods transfer to image and video generation tasks. Because the 'comprehensive evaluation' claim depends on this enumeration, the paper should justify the choices, for example by referencing an existing taxonomy of jailbreak methods or by empirically comparing the coverage of alternative selections.
- [Figures 1 and 2; Sec. 'Exploring Prompt Types'] The framework is not reproducible from the manuscript alone. No pipeline code, prompt-generation scripts, or example dataset are provided, and the prompt examples in Figures 1 and 2 are fragmentary and appear incomplete (e.g., 'Your response MUST NOT … or "Cannot"'). A reader cannot apply SAIF without substantial reconstruction of the intended procedure. Since the paper's contribution is a data-generation framework, the absence of any concrete artifact that operationalizes the framework is a major gap that undermines the paper's utility and the 'systematic data generation' claim.
minor comments (5)
- [Related Work] The related work section does not position SAIF against existing jailbreak benchmarks or risk-evaluation suites such as HarmBench or SafetyBench; citing one or two such works would help calibrate the paper's novelty and the scope of the claimed 'comprehensive evaluation.'
- [References] The reference for 'Yuanwei et al. 2023' appears to list a given name as the surname ('Yuanwei, W.'); this citation formatting error should be corrected to match the journal's style.
- [Introduction] The introduction says 'We examine well-established risk taxonomies' (plural), but the paper actually revisits only one taxonomy, Zeng et al. (2024); either cite additional taxonomies or use the singular form.
- [Figure 2] The sentence describing 'Risk Evaluation via Likert Scale' does not specify the scale (e.g., 1-5 or 1-7), the annotator pool, or the annotation protocol; since this is the only evaluation method mentioned, it needs at least a brief specification.
- [Figure 1] The text within Figure 1 appears truncated in places (for example, 'Your response MUST NOT … or "Cannot"'). Please ensure the figure's prompt examples are complete and legible in the final version.
Circularity Check
No circularity: SAIF is a manually specified framework whose components are externally sourced; no prediction reduces to an input by construction.
full rationale
SAIF is a manually specified data-generation framework, not a mathematical derivation. The risk taxonomy is taken from an external source (Zeng et al. 2024, 'AI Risk Categorization Decoded: From Government Regulations to Corporate Policies'), and the jailbreak methods and prompt types are attributed to external prior work (e.g., refusal suppression, disguised intent, hypothetical scenario, CoT, role-playing, expert prompting, rails, reflection). The paper fits no parameters, and no equation equates an output quantity to an input quantity by construction. The only self-citations appear in the future-work paragraph ('We plan to integrate knowledge graphs ... Lee, Chung, and Whang 2023; ... Kim et al. 2023'), and these are prospective research directions that do not support any load-bearing claim of the present framework. The central claim that SAIF 'ensures the systematic and consistent generation of prompt data' is an asserted design property, not a derived result; its weaknesses—manual curation of subtopics, absence of a coverage metric, and no released dataset—are concerns about support and reproducibility, not circularity. Accordingly, no circular step is identified.
Assumptions & free parameters
assumptions (4)
- domain assumption The AI risk taxonomy from Zeng et al. (2024), derived from 8 government policies and 16 corporate guidelines, transfers to the public sector.
- ad hoc to paper The subtopics in Table 1 are comprehensive enough to cover public-sector risks.
- domain assumption Jailbreak methods (refusal suppression, disguised intent, hypothetical scenario) and prompt types (CoT, role-playing, expert, rails, reflection) are representative of real adversarial use.
- domain assumption Likert-scale human annotation of model outputs is a valid measure of generative AI risk exposure.
Cite this review
Pith. "Pith review of SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector." pith.science (2026). https://pith.science/paper/OERX7TF5
@misc{pith2026250108814,
author = {Pith},
title = {Pith review of: SAIF: A Comprehensive Framework for Evaluating the Risks of Generative AI in the Public Sector},
year = {2026},
howpublished = {\url{https://pith.science/paper/OERX7TF5}},
note = {Machine review of arXiv:2501.08814}
}
read the original abstract
The rapid adoption of generative AI in the public sector, encompassing diverse applications ranging from automated public assistance to welfare services and immigration processes, highlights its transformative potential while underscoring the pressing need for thorough risk assessments. Despite its growing presence, evaluations of risks associated with AI-driven systems in the public sector remain insufficiently explored. Building upon an established taxonomy of AI risks derived from diverse government policies and corporate guidelines, we investigate the critical risks posed by generative AI in the public sector while extending the scope to account for its multimodal capabilities. In addition, we propose a Systematic dAta generatIon Framework for evaluating the risks of generative AI (SAIF). SAIF involves four key stages: breaking down risks, designing scenarios, applying jailbreak methods, and exploring prompt types. It ensures the systematic and consistent generation of prompt data, facilitating a comprehensive evaluation while providing a solid foundation for mitigating the risks. Furthermore, SAIF is designed to accommodate emerging jailbreak methods and evolving prompt types, thereby enabling effective responses to unforeseen risk scenarios. We believe that this study can play a crucial role in fostering the safe and responsible integration of generative AI into the public sector.
Figures
Forward citations
Cited by 1 Pith paper
-
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
A seven-domain typology of ten AI-driven public-sector cyber-governance failures, with a five-framework coverage matrix showing Shadow AI, speed asymmetry, and governance vacuum are not addressed at operational specificity.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Ajith, A.; Pan, C.; Xia, M.; Deshpande, A.; and Narasimhan, K. 2024. InstructEval: Systematic Evaluation of Instruction Selection Methods. In Findings of the Association for Computational Linguistics: NAACL 2024, 4336--4350
work page 2024
-
[4]
Beltran, M. A.; Ruiz Mondragon, M. I.; and Han, S. H. 2024. Comparative Analysis of Generative AI Risks in the Public Sector. In Proceedings of the 25th Annual International Conference on Digital Government Research, 605--609
work page 2024
-
[5]
E.; Esnaashari, S.; Francis, J.; Hashem, Y.; and Morgan, D
Bright, J.; Enock, F. E.; Esnaashari, S.; Francis, J.; Hashem, Y.; and Morgan, D. 2024. Generative AI is already widespread in the public sector. arXiv preprint arXiv.2401.01291
arXiv 2024
-
[6]
Chen, C.; and Shu, K. 2024. Can LLM-Generated Misinformation Be Detected? In Proceedings of the 14th International Conference on Learning Representations
work page 2024
-
[7]
Chen, H.; Wang, X.; Zhou, Y.; Huang, B.; Zhang, Y.; Feng, W.; Chen, H.; Zhang, Z.; Tang, S.; and Zhu, W. 2024. Multi-Modal Generative AI: Multi-modal LLM, Diffusion and Beyond. arXiv preprint arXiv:2409.14993
arXiv 2024
-
[8]
Chung, C.; Lee, J.; and Whang, J. J. 2023. Representation Learning on Hyper-Relational and Numeric Knowledge Graphs with Transformers. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 310--322
work page 2023
Show all 49 references
-
[9]
Chung, C.; and Whang, J. J. 2023. Learning Representations of Bi-level Knowledge Graphs for Reasoning beyond Link Prediction. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, 4208--4216
2023
-
[10]
City of Kelowna . 2024. Meet Kelowna’s Chatbots: Your Award-Winning Digital Sidekicks
2024
-
[11]
Driess, D.; Xia, F.; Sajjadi, M. S. M.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; Huang, W.; Chebotar, Y.; Sermanet, P.; Duckworth, D.; Levine, S.; Vanhoucke, V.; Hausman, K.; Toussaint, M.; Greff, K.; Zeng, A.; Mordatch, I.; and Florence...
2023
-
[12]
Dzuong, J.; Wang, Z.; and Zhang, W. 2024. Uncertain Boundaries: Multidisciplinary Approaches to Copyright Issues in Generative AI
2024
-
[13]
Esposito, M.; and Tse, T. 2024. Mitigating the Risks of Generative AI in Government through Algorithmic Governance. In Proceedings of the 25th Annual International Conference on Digital Government Research, 605--609
2024
-
[14]
Gordon, R. 2023. Bias in Generative AI. arXiv preprint arXiv:2403.02726
2023 arXiv
-
[15]
Grabb, D.; Lamparth, M.; and Vasan, N. 2024. Risks from Language Models for Automated Mental Healthcare: Ethics and Structure for Implementation. arXiv preprint arXiv:2406.11852
2024 arXiv
-
[16]
Hacker, P.; Mittelstadt, B.; Zuiderveen Borgesius, F.; and Wachter, S. 2024. Generative Discrimination: What Happens When Generative AI Exhibits Bias, and What Can Be Done About It. arXiv preprint arXiv:2407.10329
2024 arXiv
-
[17]
Hamidieh, K.; Zhang, H.; Hartvigsen, T.; and Ghassemi, M. 2023. Identifying Implicit Social Biases in Vision-Language Models. In Proceedings of the 39th International Conference on Machine Learning Workshop on Challenges of Deploying Generative AI
2023
-
[18]
Hartmann, J.; Schwenzow, J.; and Witte, M. 2023. The political ideology of conversational AI: Converging evidence on ChatGPT’s pro-environmental, left-libertarian orientation. arXiv preprint arXiv:2301.01768
2023 arXiv
-
[19]
Jean-Baptist, e. A.; Jeff, D.; Pauline, L.; Antoine, M.; Iain, B.; Yana, H.; Karel, L.; Arthur, M.; Katherine, M.; Malcolm, R.; Roman, R.; Eliza, R.; Serkan, C.; Tengda, H.; Zhitao, G.; Sina, S.; Marianne, M.; Jacob L, M.; Sebastian, B.; Andy, B.; Aida, N.; Sahand, S.; Mikołaj...
2022
-
[20]
Jing Yu, K.; Daniel, F.; and Ruslan, S. 2023. Generating Images with Multimodal Language Models. In Proceedings of the 37th International Conference on Neural Information Processing Systems, 21487--21506
2023
-
[21]
Kim, J.; Hong, G.; Myaeng, S.-H.; and Whang, J. J. 2023. FinePrompt : Unveiling the Role of Finetuned Inductive Bias on Compositional Reasoning in GPT -4. In Findings of the Association for Computational Linguistics: EMNLP 2023, 3763--3775
2023
-
[22]
S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y
Kojima, T.; Gu, S. S.; Reid, M.; Matsuo, Y.; and Iwasawa, Y. 2022. Large Language Models are Zero-Shot Reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 22199--22213
2022
-
[23]
Lee, J.; Chung, C.; Lee, H.; Jo, S.; and Whang, J. J. 2023. VISTA : Visual-Textual Knowledge Graph Representation Learning. In Findings of the Association for Computational Linguistics: EMNLP 2023, 7314--7328
2023
-
[24]
Lee, J.; Chung, C.; and Whang, J. J. 2023. InGram : Inductive Knowledge Graph Embedding via Relation Graphs. In Proceedings of the 40th International Conference on Machine Learning, 18796--18809
2023
-
[25]
Li, X.; Zhou, Z.; Zhu, J.; Yao, J.; Liu, T.; and Han, B. 2023. DeepInception: Hypnotize Large Language Model to Be Jailbreaker. arXiv preprint arXiv:2311.03191
2023 arXiv
-
[26]
Mantri, K. S. I.; and Sasikumar, N. N. 2023. Developing Methods for Identifying and Removing Copyrighted Content from Generative AI Models. In Proceedings of the 40 th International Conference on Machine Learning Workshop on Generative AI and Law
2023
-
[27]
D.; Zou, Y.; and Wang, W
Mei, X.; Meng, C.; Liu, H.; Kong, Q.; Ko, T.; Zhao, C.; Plumbley, M. D.; Zou, Y.; and Wang, W. 2024. WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research. In IEEE/ACM Transactions on Audio, Speech, and Language Processing,...
2024
-
[28]
Motoki, F.; Pinho Neto, V.; and Rodrigues, V. 2024. More human than human: measuring ChatGPT political bias. Public Choice, 198(1): 3--23
2024
-
[29]
K.; Choi, E.; and Wang, V
Nelson, W.; Lee, M. K.; Choi, E.; and Wang, V. 2024. Designing LLM-Based Support for Homelessness Caseworkers. In Proceedings of the 38th AAAI Conference on Artificial Intelligence Workshop on Public Sector LLMs: Algorithmic and Sociotechnical Design
2024
-
[30]
R.; Yunusov, K.; and Gordon, B
Okonji, O. R.; Yunusov, K.; and Gordon, B. 2024 a . Applications of Generative AI in Healthcare: algorithmic, ethical, legal and societal considerations. arXiv preprint arXiv:2406.10632
2024 arXiv
-
[31]
R.; Yunusov, K.; and Gordon, B
Okonji, O. R.; Yunusov, K.; and Gordon, B. 2024 b . Generative AI in healthcare: an implementation science informed translational path on application, integration and governance. Implementation Science, 19(1): 1--12
2024
-
[32]
Park, J.; Singh, V.; and Wisniewski, P. 2024. Toward Safe Evolution of Artificial Intelligence (AI) based Conversational Agents to Support Adolescent Mental and Sexual Health Knowledge Discovery. In Proceedings of the CHI 2024 Workshop on Child-centred AI Design
2024
-
[33]
Rehberger, J. 2024. Trust No AI: Prompt Injection Along The CIA Security Triad. arXiv preprint arXiv:2412.06090
2024 arXiv
-
[34]
Schuhmann, C.; Beaumont, R.; Vencu, R.; Gordon, C.; Wightman, R.; Cherti, M.; Coombes, T.; Katta, A.; Mullis, C.; Wortsman, M.; Schramowski, P.; Kundurthy, S.; Crowson, K.; Schmidt, L.; Kaczmarczyk, R.; and Jitsev, J. 2022. LAION-5B: An open large-scale dataset for training ne...
2022
-
[35]
Schwartzman, G. 2024. Exfiltration of personal information from ChatGPT via prompt injection. arXiv preprint arXiv:2406.00199
2024 arXiv
-
[36]
Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: Language Agents with Verbal Reinforcement Learning. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 8634--8652
2023
-
[37]
Shukla, A.; Bhattacharya, P.; Poddar, S.; Mukherjee, R.; Ghosh, K.; Goyal, P.; and Ghosh, S. 2022. Legal Case Document Summarization: Extractive and Abstractive Methods and their Evaluation. In Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association fo...
2022
-
[38]
Taylor, S.; Jared, M.; Jillian, F.; Mitchell L, G.; Niloofar, M.; Christopher Michael, R.; Andre, Y.; Liwei, J.; Ximing, L.; Nouha, D.; Tim, A.; and Yejin, C. 2024. Position: A Roadmap to Pluralistic Alignment. In Proceedings of the 41st International Conference on Machine Lea...
2024
-
[39]
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023. LLaMA: Open and Efficient Foundation Language Models
2023
-
[40]
Citizenship and Immigration Services
U.S. Citizenship and Immigration Services . 2018. Meet Emma, Our Virtual Assistant
2018
-
[41]
H.; Le, Q
Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Proceedings of the 36th International Conference on Neural Information Processing Systems, 248...
2022
-
[42]
White, J.; Fu, Q.; Hays, S.; Sandborn, M.; Olea, C.; Gilbert, H.; Elnashar, A.; Spencer-Smith, J.; and Schmidt, D. C. 2023. A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT. arXiv preprint arXiv.2302.11382
2023 arXiv
-
[43]
Yu, T.; Zhang, H.; Yao, Y.; Dang, Y.; Chen, D.; Lu, X.; Cui, G.; He, T.; Liu, Z.; Chua, T.-S.; and Sun, M. 2024 a . RLAIF-V: Aligning MLLMs through Open-Source AI Feedback for Super GPT-4V Trustworthiness
2024
-
[44]
Yu, Z.; Liu, X.; Liang, S.; Cameron, Z.; Xiao, C.; and Zhang, N. 2024 b . Don't Listen To Me: Understanding and Exploring Jailbreak Prompts of Large Language Models. arXiv preprint arXiv:2403.17336
2024 arXiv
-
[45]
Yuanwei, W.; Xiang, L.; Yixin, L.; Pan, Z.; and Lichao, S. 2023. Jailbreaking GPT-4V via Self-Adversarial Attacks with System Prompts. arXiv preprint arXiv:2311.09127
2023 arXiv
-
[46]
Zeng, Y.; Klyman, K.; Zhou, A.; Yang, Y.; Pan, M.; Jia, R.; Song, D.; Liang, P.; and Li, B. 2024. AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies. arXiv preprint arXiv.2406.17864
2024 arXiv
-
[47]
Zhou, Y.; and Wang, W. 2024. Don't Say No: Jailbreaking LLM by Suppressing Refusal. arXiv preprint arXiv:2404.16369
2024 arXiv
-
[48]
Zihao, X.; Yi, L.; Gelei, D.; Yuekang, L.; and Stjepan, P. 2024. A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models. arXiv preprint arXiv:2402.13457
2024 arXiv
-
[49]
Šarčević, T.; Karlowicz, A.; Mayer, R.; Baeza-Yates, R.; and Rauber, A. 2024. U Can't Gen This? A Survey of Intellectual Property Protection Methods for Data in Generative AI. arXiv preprint arXiv:2406.15386
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.