A protocol-level identifiability audit over finite behavioral policy classes shows that base-only observation under-identifies the selective-response estimand, full support identifies it, and base accuracy diverges from intervention-response fidelity in two instruction-tuned models.
InAdvances in Neu- ral Information Processing Systems 36 (NeurIPS 2023)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Beyond Local Accuracy: A Protocol-Level Identifiability Audit for Controlled LLM Reasoning Evaluation
A protocol-level identifiability audit over finite behavioral policy classes shows that base-only observation under-identifies the selective-response estimand, full support identifies it, and base accuracy diverges from intervention-response fidelity in two instruction-tuned models.