Pith. sign in

REVIEW 2 cited by

LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.02987 v2 pith:GNGQ6PHI submitted 2024-07-03 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords contentlora-guardmoderationguardraillanguagellmsmodelsadaptation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Guardrails have emerged as an alternative to safety alignment for content moderation of large language models (LLMs). Existing model-based guardrails have not been designed for resource-constrained computational portable devices, such as mobile phones, more and more of which are running LLM-based applications locally. We introduce LoRA-Guard, a parameter-efficient guardrail adaptation method that relies on knowledge sharing between LLMs and guardrail models. LoRA-Guard extracts language features from the LLMs and adapts them for the content moderation task using low-rank adapters, while a dual-path design prevents any performance degradation on the generative task. We show that LoRA-Guard outperforms existing approaches with 100-1000x lower parameter overhead while maintaining accuracy, enabling on-device content moderation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging the Safety Gap: A Guardrail Pipeline for Trustworthy LLM Inferences

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A guardrail pipeline combining detection, retrieval grounding, rule-based wrappers, and a repair model is reported to match OpenAI moderation and fix 80.7 percent of hallucinated HaluEval answers.

  2. Latenrgy: Model Agnostic Latency and Energy Consumption Prediction for Binary Classifiers

    cs.LG 2024-12 reject novelty 3.0 of 10

    A theoretical framework proposes two linear equations for predicting inference latency and energy across binary classifiers, but no coefficients are calibrated and no validation is provided.

Pith tools