REVIEW 1 major objections 1 minor 7 references
One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline
T0 review · 1 major / 1 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read A zero-parameter compression baseline reaches 74.7 percent weighted accuracy on all 102 Tuebingen pairs when every method must decide without tuning or abstention.
desk verdict Under one forced-decision protocol on all 102 pairs the methods land in the low-to-mid 70s and a zero-parameter compressor matches the strongest of them. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The same-hands re-evaluation protocol that applies one strict rule—no tuning and a decision forced on every pair—to all methods on the full set of 102 Tuebingen pairs, benchmarked against a sorted-conditional compression baseline that feeds quantized, sorted, first-differenced data to an off-the-shelf bz2 compressor.
What would settle it
A re-run of the same 102 pairs under the forced-decision protocol in which any literature method exceeds the baseline by a margin larger than the McNemar test noise level would falsify the reported clustering and tie.
Extended reading notes
Core claim
Under the common ruler of evaluating every method on the identical 102 pairs with forced decisions and no tuning, the sorted-conditional compression baseline reaches 74.7 percent weighted accuracy. A faithful re-run of RECI lands at 70.7 percent. SLOPE's published 82.4 percent is reproduced only when scoring is restricted to the pairs its significance test chooses to answer; on the full set the figure drops. The methods therefore cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them.
Load-bearing premise
The claim rests on the premise that forcing a binary decision on every pair without allowing significance-based abstention or model selection constitutes the correct common ruler for comparison.
Editorial extensions
If this is right
- SLOPE's published 82.4 percent reflects performance only on the subset its significance test elects to answer rather than on the full set.
- RECI's re-run score of 70.7 percent falls inside the original authors' reported error bar, not the 77.5 percent figure often quoted.
- Compression score magnitude functions as a model-free indicator of confounding (p = 2.8e-68).
- A pre-registered falsification test fails in a manner that bounds the theoretical interpretation of the compression approach.
- Under the uniform protocol all examined methods perform in the low-to-mid 70s.
Reading between the lines
- Adopting the forced-decision protocol more widely would require future causal discovery papers to report both full-set and selective accuracies for direct comparison.
- The observed clustering implies that further gains from increasingly complex models may be small once protocol differences are eliminated, shifting attention to data characteristics.
- The same-hands approach could be applied to other bivariate or low-dimensional causal benchmarks to test whether the Tuebingen convergence generalizes.
- A compression baseline of this form supplies an immediate, training-free reference point for any new method proposed on similar pair data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that published accuracies for bivariate causal direction methods on the Tuebingen pairs are not comparable due to differing protocols (pair subsets, weightings, model selection, and abstention rates). It conducts a same-hands re-evaluation forcing binary decisions on the identical 102 pairs with no tuning, introduces a zero-parameter sorted-conditional compression baseline (quantized, sorted, first-differenced data fed to bz2), and reports that this baseline reaches 74.7% weighted accuracy (p=3.7e-7) while methods cluster in the low-to-mid 70s. It reproduces lower figures for RECI (70.7%) and forced-decision SLOPE (77.2%), attributes higher published numbers to test-set selection and significance-gated abstention, and releases code, pre-registrations, and per-pair outputs.
Significance. If the re-implementations and protocol hold, the work supplies a transparent, fully reproducible parameter-free baseline together with independent statistical tests (McNemar, p-values) and a confounding flag via compression scores (p=2.8e-68). The public release of code, pre-registrations, and per-pair outputs is a clear strength that enables direct verification and improves benchmarking standards in causal discovery.
major comments (1)
- [Abstract] Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler.
minor comments (1)
- [Abstract] Abstract: The brief description of the baseline ('quantized, sorted, first-differenced data') leaves the exact quantization scheme and differencing order implicit; a one-sentence expansion would improve standalone readability even though code is released.
Simulated Author's Rebuttal
We thank the referee for the constructive comment regarding the abstract and the choice of protocol. We address the point directly below.
read point-by-point responses
-
Referee: [Abstract] Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler.
Authors: We agree that the forced binary-decision protocol is the foundation of our comparability claim and that it necessarily lowers SLOPE from its published decided-subset figure (which we already reproduce at 81.7 %). Our manuscript explicitly contrasts the two figures and attributes the difference to significance-gated abstention. Nevertheless, the referee's request for an explicit sensitivity analysis is reasonable. In the revised manuscript we will add a dedicated subsection that (i) scores all abstentions as errors (zero contribution to accuracy) for every method that permits them and (ii) applies proper scoring rules (Brier score and log-loss) to any probabilistic outputs that are available. This will allow readers to verify whether the low-to-mid-70s cluster persists under alternative treatments of abstention. revision: yes
Circularity Check
No significant circularity; baseline and evaluations are independent of fitted parameters or self-citation chains.
full rationale
The paper's central claims rest on a parameter-free compression baseline (sorted-conditional compression using bz2 on quantized, sorted, first-differenced data) applied to the external Tuebingen dataset under a fixed protocol of forced binary decisions on all 102 pairs. Accuracies (e.g., 74.7% weighted) and p-values (e.g., 3.7e-7) are computed directly from these runs and standard statistical tests (McNemar) without any parameter fitting, model selection on the test set, or reduction to quantities defined by the authors' prior work. No self-citations are load-bearing for the core result; the re-evaluations of other methods (SLOPE, RECI) reproduce published outputs or stored decisions on the same fixed pairs. The derivation chain is self-contained against external benchmarks and does not exhibit self-definitional, fitted-input, or ansatz-smuggling patterns.
Assumptions & free parameters
assumptions (2)
- domain assumption The Tuebingen cause-effect pairs constitute a suitable fixed benchmark for comparing bivariate causal direction methods.
- ad hoc to paper Forcing a decision on every pair without abstention produces a fairer comparison than significance-gated protocols.
Cite this review
Pith. "Pith review of One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline." pith.science (2026). https://pith.science/paper/DK6PCOQP
@misc{pith2026260623767,
author = {Pith},
title = {Pith review of: One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline},
year = {2026},
howpublished = {\url{https://pith.science/paper/DK6PCOQP}},
note = {Machine review of arXiv:2606.23767}
}
read the original abstract
Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors' own protocol -- different pair subsets, weightings, model-selection, and decision rates. We argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one strict rule -- no tuning and a decision forced on every pair. As a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data to an off-the-shelf compressor (bz2) and has zero fitted parameters. Under the common ruler the ranking differs sharply from the literature. Our baseline reaches 74.7% weighted accuracy (p = 3.7e-7); on the same 100 pairs that SLOPE is evaluated on it scores 76.0%, a 1.2-point gap below the authors' own forced-decision SLOPE (77.2%) that is well inside noise (McNemar p = 0.39). A faithful re-run of RECI lands at 70.7% -- inside the original authors' reported error bar, not the 77.5% often quoted (which we trace to a mis-copied cell). SLOPE's published 82.4% is a decided-subset figure: scoring the authors' own stored output only on the pairs its significance test chose to answer reproduces 81.7%. Under the common ruler the methods cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them. We document the mechanisms that inflate published figures (test-set model selection, significance-gated abstention) and contribute two further results: compression score magnitude is a model-free confounding flag (p = 2.8e-68), and a pre-registered falsification test fails in an instructive way that bounds the method's theoretical interpretation. Code, pre-registrations, and per-pair outputs are released.
Reference graph
Works this paper leans on
-
[1]
Mooij, J
J. Mooij, J. Peters, D. Janzing, J. Zscheischler, B. Schölkopf. Distinguishing cause from effect using observational data: methods and benchmarks.JMLR17(32):1–102, 2016
2016
-
[2]
Janzing, B
D. Janzing, B. Schölkopf. Causal inference using the algorith- mic Markov condition.IEEE Trans. Inf. Theory56(10):5168– 5194, 2010
2010
-
[3]
Lemeire, D
J. Lemeire, D. Janzing. Replacing causal faithfulness with algorithmic independence.Minds and Machines23(2):227– 249, 2013
2013
-
[4]
Janzing et al
D. Janzing et al. Information-geometric approach to inferring causal directions.Artificial Intelligence182–183:1–31, 2012
2012
-
[5]
Blöbaum, D
P. Blöbaum, D. Janzing, T. Washio, S. Shimizu, B. Schölkopf. Cause-effect inference by comparing regression errors.AIS- TATS2018; extended analysis,PeerJ CS2019. 5
-
[6]
A. Marx, J. Vreeken. Telling cause from effect using MDL- based local and global regression.ICDM2017
-
[7]
K. Hlavackova-Schindler, A. Marsela. Identifying causal di- rection via dense functional classes. arXiv:2509.00538, 2025. 6
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.