Introduces CounterBench benchmark and CoIn iterative reasoning method showing LLMs perform near random on formal counterfactual tasks but improve substantially with guided backtracking.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
representative citing papers
Prediction-tuned synthetic data can pass realism tests while distorting average treatment effects, and separating covariate generation from treatment/outcome modeling largely fixes it.
citing papers explorer
-
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
Introduces CounterBench benchmark and CoIn iterative reasoning method showing LLMs perform near random on formal counterfactual tasks but improve substantially with guided backtracking.
-
Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
Prediction-tuned synthetic data can pass realism tests while distorting average treatment effects, and separating covariate generation from treatment/outcome modeling largely fixes it.