REVIEW 14 cited by
A Safe Harbor for AI Evaluation and Red Teaming
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Independent evaluation and red teaming are critical for identifying the risks posed by generative AI systems. However, the terms of service and enforcement strategies used by prominent AI companies to deter model misuse have disincentives on good faith safety evaluations. This causes some researchers to fear that conducting such research or releasing their findings will result in account suspensions or legal reprisal. Although some companies offer researcher access programs, they are an inadequate substitute for independent research access, as they have limited community representation, receive inadequate funding, and lack independence from corporate incentives. We propose that major AI developers commit to providing a legal and technical safe harbor, indemnifying public interest safety research and protecting it from the threat of account suspensions or legal reprisal. These proposals emerged from our collective experience conducting safety, privacy, and trustworthiness research on generative AI systems, where norms and incentives could be better aligned with public interests, without exacerbating model misuse. We believe these commitments are a necessary step towards more inclusive and unimpeded community efforts to tackle the risks of generative AI.
Forward citations
Cited by 14 Pith papers
-
Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
For demographic parity and binary cross-entropy loss, the paper gives a closed-form upper bound for the loss-fairness Pareto frontier, used as a scaling law to audit large models without training them.
-
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Legal and ethical bans on CSAM access and generation break standard AI safety techniques, creating 15 open problems that demand new methods for dataset cleaning, concept fusion prevention, fine-tuning resilience, dete...
-
Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments
A TEE-based protocol for cryptographically verifiable AI safety benchmark results, demonstrated on Llama-3.1 with AWS Nitro Enclaves.
-
Adversarial Attacks on Robotic Vision Language Action Models
Text-based adversarial suffixes can make OpenVLA robot policies elicit chosen target actions with over 90% success on one-hot targets and persist across rollout steps.
-
Real-World Gaps in AI Governance Research
Corporate AI safety research is dominated by pre-deployment alignment and evaluation work, while high-risk deployment topics such as medical error, misinformation, bias, behavioral design, and copyright are measured t...
-
When Testing AI Tests Us: Safeguarding Mental Health on the Digital Frontlines
AI red-teaming can harm the mental health of the people who do it, and protective practices from four comparable professions can be adapted to support them.
-
Access Denied: Meaningful Data Access for Quantitative Algorithm Audits
Simulated algorithm audits show that synthetic data and small or incomplete audit samples can make group-parity metrics unreliable, while differentially private aggregate statistics generally remain reliable.
-
WeAudit: Scaffolding User Auditors and AI Practitioners in Auditing Generative AI
WeAudit scaffolds everyday users to audit generative AI through comparison, examples, discussion, verification, and structured reports, and practitioners found the resulting audit reports actionable.
-
A Framework for Evaluating LLMs Under Task Indeterminacy
When evaluation items admit multiple valid responses, gold-label accuracy underestimates true model performance, and the paper offers bounds on the true performance from partial knowledge.
-
The Pitfalls of "Security by Obscurity" And What They Mean for Transparent AI
Security's hard-won transparency practices, from Kerckhoffs' principle to vulnerability disclosure, form three transferable themes for AI transparency efforts.
-
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
GuardVal combines role-playing jailbreak generation with an Adam-inspired optimizer and an Overall Safety Value metric, but the method is underspecified and not validated with released code or data.
-
Position Paper: Model Access should be a Key Concern in AI Governance
Model access decisions should be studied and coordinated through a dedicated research field, with recommendations for evaluators, companies, governments, and international bodies.
-
AI Safety Frameworks Should Include Procedures for Model Access Decisions
The paper introduces Responsible Access Policies as a recommended component of frontier AI safety frameworks for governing model access.
-
Enabling External Scrutiny of AI Systems with Privacy-Enhancing Technologies
OpenMined's privacy-enhancing infrastructure enabled two pilot AI audits, and the authors argue the remaining barriers to routine external scrutiny are legal rather than technical.
Discussion (0). Continue with ORCID to comment.