Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.
Abstracting Causal Models
3 Pith papers cite this work. Polarity classification is still indexing.
abstract
We consider a sequence of successively more restrictive definitions of abstraction for causal models, starting with a notion introduced by Rubenstein et al. (2017) called exact transformation that applies to probabilistic causal models, moving to a notion of uniform transformation that applies to deterministic causal models and does not allow differences to be hidden by the "right" choice of distribution, and then to abstraction, where the interventions of interest are determined by the map from low-level states to high-level states, and strong abstraction, which takes more seriously all potential interventions in a model, not just the allowed interventions. We show that procedures for combining micro-variables into macro-variables are instances of our notion of strong abstraction, as are all the examples considered by Rubenstein et al.
fields
cs.AI 3representative citing papers
Derivation graphs characterize the space of do-calculus equivalent interventional expressions, enable identification with at most four rule applications, and yield multiple valid estimands for improved efficiency.
Extends exact causal abstraction to approximate abstractions for causal models, including probabilistic versions, to handle discrepancies between abstraction levels.
citing papers explorer
-
Safety from Honesty in a Disinterested AI Predictor
Under consequence-invariant posterior training and sparsity of coordinated harm patterns, the training mass on dangerous guarded Predictors is bounded by C_bad times R_shell.
-
Unveiling the Structure of Do-Calculus Reasoning via Derivation Graphs
Derivation graphs characterize the space of do-calculus equivalent interventional expressions, enable identification with at most four rule applications, and yield multiple valid estimands for improved efficiency.
-
Approximate Causal Abstraction
Extends exact causal abstraction to approximate abstractions for causal models, including probabilistic versions, to handle discrepancies between abstraction levels.