REVIEW 8 cited by
CausalBench: A Large-scale Benchmark for Network Inference from Single-cell Perturbation Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Causal inference is a vital aspect of multiple scientific disciplines and is routinely applied to high-impact applications such as medicine. However, evaluating the performance of causal inference methods in real-world environments is challenging due to the need for observations under both interventional and control conditions. Traditional evaluations conducted on synthetic datasets do not reflect the performance in real-world systems. To address this, we introduce CausalBench, a benchmark suite for evaluating network inference methods on real-world interventional data from large-scale single-cell perturbation experiments. CausalBench incorporates biologically-motivated performance metrics, including new distribution-based interventional metrics. A systematic evaluation of state-of-the-art causal inference methods using our CausalBench suite highlights how poor scalability of current methods limits performance. Moreover, methods that use interventional information do not outperform those that only use observational data, contrary to what is observed on synthetic benchmarks. Thus, CausalBench opens new avenues in causal network inference research and provides a principled and reliable way to track progress in leveraging real-world interventional data.
Forward citations
Cited by 8 Pith papers
-
R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation
A 3,068-prompt benchmark with per-instance Q&A scoring shows that current text-to-image models, including reasoning-enhanced ones, handle reasoning-driven prompts poorly, with mathematical reasoning near zero.
-
Causal ASCEND: Scalable Two-tier Causal Discovery on High Dimensional Multi-omics Data
ASCEND recovers ancestral gene relationships at omics scale by conditioning only on dynamically updated nearest ancestors under a known two-tier ordering, with polynomial complexity and higher precision than GRN and c...
-
Causality-Inspired Robustness for Nonlinear Models via Representation Learning
The proposed two-step method, CIRRL, learns an affine-equivalent latent representation of causal structure and applies a distributionally robust linear regression on it, claiming minimax robustness against bounded dis...
-
Achievable distributional robustness when the robust risk is only partially identified
In partially identified linear models, the worst-case robust risk yields a minimax predictor that abstains on unseen shift directions and provably outperforms anchor regression and OLS when test shifts include new directions.
-
GPO-VAE: Modeling Explainable Gene Perturbation Responses utilizing GRN-Aligned Parameter Optimization
GPO-VAE reinterprets a VAE's perturbation parameters as a gene-by-gene causal matrix, trains it with a differential-expression-matching loss, and reports state-of-the-art perturbation prediction plus GRN inference on ...
-
Multi-megabase scale genome interpretation with genetic language models
Phenformer uses frozen Enformer embeddings of 512 gene windows to predict disease risk and cell type involvement, outperforming PRS baselines restricted to the same genomic regions.
-
Scalable Temporal Anomaly Causality Discovery in Large Systems: Achieving Computational Efficiency with Binary Anomaly Flag Data
Temporal causal graphs can be learned from binary alarm flags with 99% data compression and modest accuracy gains using heuristic modifications to PCMCI.
-
The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications
A systematic review showing causal discovery evaluation is still dominated by small, low-diversity datasets and structural metrics, with a curated set of realistic alternatives.
Discussion (0). Continue with ORCID to comment.