Pith. sign in

Detectors for Safe and Reliable LLMs: Implementations, Uses, and Limitations

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Large language models (LLMs) are susceptible to a variety of risks, from non-faithful output to biased and toxic generations. Due to several limiting factors surrounding LLMs (training cost, API access, data availability, etc.), it may not always be feasible to impose direct safety constraints on a deployed model. Therefore, an efficient and reliable alternative is required. To this end, we present our ongoing efforts to create and deploy a library of detectors: compact and easy-to-build classification models that provide labels for various harms. In addition to the detectors themselves, we discuss a wide range of uses for these detector models - from acting as guardrails to enabling effective AI governance. We also deep dive into inherent challenges in their development and discuss future work aimed at making the detectors more reliable and broadening their scope.

fields

cs.LG 1

years

2025 1

verdicts

REJECT 1

representative citing papers

Paying Alignment Tax with Contrastive Learning

cs.LG · 2025-05-25 · reject · novelty 5.0

A contrastive learning framework with positive and negative example pairs improves faithfulness and slightly reduces toxicity on Reddit TL;DR summarization, but the central claim of avoiding the alignment tax is not evaluated on the paper's own knowledge benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Paying Alignment Tax with Contrastive Learning cs.LG · 2025-05-25 · reject · none · ref 3 · internal anchor

    A contrastive learning framework with positive and negative example pairs improves faithfulness and slightly reduces toxicity on Reddit TL;DR summarization, but the central claim of avoiding the alignment tax is not evaluated on the paper's own knowledge benchmarks.