Pith. sign in

REVIEW 14 cited by

garak: A Framework for Security Probing Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11036 v1 pith:OMDURBYW submitted 2024-06-16 cs.CL cs.CR

classification cs.CLcs.CR
keywords securitymodelsframeworkgaraklanguagetargetvulnerabilitieswhat
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As Large Language Models (LLMs) are deployed and integrated into thousands of applications, the need for scalable evaluation of how models respond to adversarial attacks grows rapidly. However, LLM security is a moving target: models produce unpredictable output, are constantly updated, and the potential adversary is highly diverse: anyone with access to the internet and a decent command of natural language. Further, what constitutes a security weak in one context may not be an issue in a different context; one-fits-all guardrails remain theoretical. In this paper, we argue that it is time to rethink what constitutes ``LLM security'', and pursue a holistic approach to LLM security evaluation, where exploration and discovery of issues are central. To this end, this paper introduces garak (Generative AI Red-teaming and Assessment Kit), a framework which can be used to discover and identify vulnerabilities in a target LLM or dialog system. garak probes an LLM in a structured fashion to discover potential vulnerabilities. The outputs of the framework describe a target model's weaknesses, contribute to an informed discussion of what composes vulnerabilities in unique contexts, and can inform alignment and policy discussions for LLM deployment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection

    cs.CR 2026-04 unverdicted novelty 7.0 of 10

    The work introduces and partially evaluates seven cross-domain prompt injection detectors, reporting F1 gains on benchmarks like deepset/prompt-injections and indirect-injection sets via local alignment, stylometry, a...

  2. AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation

    cs.CR 2026-07 conditional novelty 6.0 of 10

    A phase-structured multi-turn red-team framework reports 97.6–100% lenient ASR but only 66.7–78.6% full actionable ASR on six frontier LLMs, with success strongly depth-dependent.

  3. LLM in the Middle: A Systematic Review of Threats and Mitigations to Real-World LLM-based Systems

    cs.CR 2025-09 conditional novelty 6.0 of 10

    A systematic review that categorizes LLM threats, severity scores, and mitigations across development and operation life cycles and multiple deployment scenarios.

  4. Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    Reasoning-augmented LLMs are on average about 3 points more robust to prompt attacks, but category-level results flip this, including a 32-point higher success rate for tree-of-attacks jailbreaks.

  5. Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset

    cs.CR 2025-05 conditional novelty 6.0 of 10

    An evaluation of seven LLM security tools on a new 500-prompt benchmark finds the ChatGPT-3.5-Turbo baseline unusable due to false positives and names Lakera Guard and ProtectAI LLM Guard the best overall tools.

  6. Importing Phantoms: Measuring LLM Package Hallucination Vulnerabilities

    cs.LG 2025-01 conditional novelty 6.0 of 10

    LLMs hallucinate non-existent software packages at rates up to 46%, and larger models with higher HumanEval scores show lower hallucination rates.

  7. A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    Introduces a multi-role red teaming framework using attacker and jury models that increases attack success rates by up to 7.9% on LLM faithfulness in question-answering tasks.

  8. aiXamine: Simplified LLM Safety and Security

    cs.CR 2025-04 conditional novelty 5.0 of 10

    The paper presents aiXamine, a black-box LLM safety and security evaluation platform that aggregates 40+ existing benchmarks into 8 services, and reports a leaderboard of 16 models showing specific vulnerabilities in ...

  9. OneShield -- the Next Generation of LLM Guardrails

    cs.CR 2025-07 conditional novelty 4.0 of 10

    A paper describes OneShield, a model-agnostic guardrail framework with parallel risk detectors and a policy manager, and reports its enterprise deployment and use in InstructLab.

  10. Measuring What Matters: A Framework for Evaluating Safety Risks in Real-World LLM Applications

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A practical framework for application-level LLM safety testing: organization-specific taxonomies plus black-box adversarial evaluation, illustrated by a Singapore government pilot.

  11. JavelinGuard: Low-Cost Transformer Architectures for LLM Security

    cs.LG 2025-06 reject novelty 4.0 of 10

    A study of five small transformer classifier architectures for LLM jailbreak and prompt injection detection claims low-latency accuracy comparable to large models, led by the multi-task Raudra design.

  12. Improved Large Language Model Jailbreak Detection via Pretrained Embeddings

    cs.CR 2024-12 reject novelty 4.0 of 10

    A random forest on Snowflake embeddings detects jailbreak prompts with F1 0.96 on JailbreakHub, but the high score depends on training on the same in-the-wild jailbreak source used in that benchmark.

  13. Generating Attacks for LLMs with GFlowNets

    cs.AI 2026-08 conditional novelty 3.0 of 10

    The authors apply GFlowNet-based reinforcement learning to train LLMs that generate English and Turkish adversarial prompts, reporting improved red-teaming success rates over a prior English-only method.

  14. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

Pith tools