REVIEW 4 cited by
Assessing Language Model Deployment with Risk Cards
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This paper introduces RiskCards, a framework for structured assessment and documentation of risks associated with an application of language models. As with all language, text generated by language models can be harmful, or used to bring about harm. Automating language generation adds both an element of scale and also more subtle or emergent undesirable tendencies to the generated text. Prior work establishes a wide variety of language model harms to many different actors: existing taxonomies identify categories of harms posed by language models; benchmarks establish automated tests of these harms; and documentation standards for models, tasks and datasets encourage transparent reporting. However, there is no risk-centric framework for documenting the complexity of a landscape in which some risks are shared across models and contexts, while others are specific, and where certain conditions may be required for risks to manifest as harms. RiskCards address this methodological gap by providing a generic framework for assessing the use of a given language model in a given scenario. Each RiskCard makes clear the routes for the risk to manifest harm, their placement in harm taxonomies, and example prompt-output pairs. While RiskCards are designed to be open-source, dynamic and participatory, we present a "starter set" of RiskCards taken from a broad literature survey, each of which details a concrete risk presentation. Language model RiskCards initiate a community knowledge base which permits the mapping of risks and harms to a specific model or its application scenario, ultimately contributing to a better, safer and shared understanding of the risk landscape.
Forward citations
Cited by 4 Pith papers
-
Oversight Structures for Agentic AI in Public-Sector Organizations
Agentic AI will intensify three existing public-sector governance challenges: continuous oversight, integrating governance with operations, and cross-departmental coordination.
-
Understanding Gender Bias in AI-Generated Product Descriptions
AI-generated product descriptions on eBay show systematic gender bias, including body-size exclusions, stereotyped feature emphasis, and differences in calls to action.
-
Developing a Risk Identification Framework for Foundation Model Uses
The paper derives four design requirements for use-based foundation model risk identification and presents an initial questionnaire-based framework demonstrated on a visitor-center chatbot example.
-
OneShield -- the Next Generation of LLM Guardrails
A paper describes OneShield, a model-agnostic guardrail framework with parallel risk detectors and a policy manager, and reports its enterprise deployment and use in InstructLab.
Discussion (0). Sign in to comment.