Does this approach strongly depend on handcrafted features, expert supervision, or human reliability? □

Human Unreliability

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

browse 2 citing papers

representative citing papers

The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning

cs.LG · 2024-03-05 · unverdicted · novelty 6.0

WMDP is a public benchmark measuring hazardous LLM knowledge across biosecurity, cybersecurity, and chemical security, paired with RMU unlearning that reduces WMDP performance without degrading general capabilities.

Representation Engineering: A Top-Down Approach to AI Transparency

cs.LG · 2023-10-02 · unverdicted · novelty 6.0

Representation engineering uses population-level representations in deep neural networks to monitor and manipulate cognitive phenomena like honesty and harmlessness, providing simple effective baselines for LLM safety.

citing papers explorer

Showing 2 of 2 citing papers.

The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning cs.LG · 2024-03-05 · unverdicted · none · ref 8
WMDP is a public benchmark measuring hazardous LLM knowledge across biosecurity, cybersecurity, and chemical security, paired with RMU unlearning that reduces WMDP performance without degrading general capabilities.
Representation Engineering: A Top-Down Approach to AI Transparency cs.LG · 2023-10-02 · unverdicted · none · ref 15
Representation engineering uses population-level representations in deep neural networks to monitor and manipulate cognitive phenomena like honesty and harmlessness, providing simple effective baselines for LLM safety.

Does this approach strongly depend on handcrafted features, expert supervision, or human reliability? □

fields

years

verdicts

representative citing papers

citing papers explorer