Pith. sign in

REVIEW 1 major objections 1 minor 7 references

One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline

T0 review · 1 major / 1 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read A zero-parameter compression baseline reaches 74.7 percent weighted accuracy on all 102 Tuebingen pairs when every method must decide without tuning or abstention.

desk verdict Under one forced-decision protocol on all 102 pairs the methods land in the low-to-mid 70s and a zero-parameter compressor matches the strongest of them. read the letter →

arxiv 2606.23767 v1 pith:DK6PCOQP submitted 2026-06-22 cs.LG

classification cs.LG
keywords causaldiscoveryTuebingencause-effectpairsbivariatedirectioncompressionbaselinere-evaluationprotocolparameter-freemethodforceddecisionevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper establishes that routine comparisons of bivariate causal direction methods on the Tuebingen cause-effect pairs rest on inconsistent protocols that differ in pair subsets, weightings, model selection, and decision rates. It therefore applies one uniform protocol: every method is re-run on the identical 102 pairs with no tuning permitted and a binary decision required for every pair. Under this common ruler a deliberately minimal baseline that sorts, quantizes, first-differences the data and feeds it to an off-the-shelf bz2 compressor scores 74.7 percent weighted accuracy. Re-evaluations of published methods show that their higher headline numbers often arise from scoring only on decided subsets or from test-set model selection. The result is that accuracies cluster tightly in the low-to-mid 70s, with the parameter-free compressor tying the strongest competitors.

What carries the argument

The same-hands re-evaluation protocol that applies one strict rule—no tuning and a decision forced on every pair—to all methods on the full set of 102 Tuebingen pairs, benchmarked against a sorted-conditional compression baseline that feeds quantized, sorted, first-differenced data to an off-the-shelf bz2 compressor.

What would settle it

A re-run of the same 102 pairs under the forced-decision protocol in which any literature method exceeds the baseline by a margin larger than the McNemar test noise level would falsify the reported clustering and tie.

Watch

Extended reading notes

Core claim

Under the common ruler of evaluating every method on the identical 102 pairs with forced decisions and no tuning, the sorted-conditional compression baseline reaches 74.7 percent weighted accuracy. A faithful re-run of RECI lands at 70.7 percent. SLOPE's published 82.4 percent is reproduced only when scoring is restricted to the pairs its significance test chooses to answer; on the full set the figure drops. The methods therefore cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them.

Load-bearing premise

The claim rests on the premise that forcing a binary decision on every pair without allowing significance-based abstention or model selection constitutes the correct common ruler for comparison.

Editorial extensions

If this is right

  • SLOPE's published 82.4 percent reflects performance only on the subset its significance test elects to answer rather than on the full set.
  • RECI's re-run score of 70.7 percent falls inside the original authors' reported error bar, not the 77.5 percent figure often quoted.
  • Compression score magnitude functions as a model-free indicator of confounding (p = 2.8e-68).
  • A pre-registered falsification test fails in a manner that bounds the theoretical interpretation of the compression approach.
  • Under the uniform protocol all examined methods perform in the low-to-mid 70s.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Adopting the forced-decision protocol more widely would require future causal discovery papers to report both full-set and selective accuracies for direct comparison.
  • The observed clustering implies that further gains from increasingly complex models may be small once protocol differences are eliminated, shifting attention to data characteristics.
  • The same-hands approach could be applied to other bivariate or low-dimensional causal benchmarks to test whether the Tuebingen convergence generalizes.
  • A compression baseline of this form supplies an immediate, training-free reference point for any new method proposed on similar pair data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The manuscript argues that published accuracies for bivariate causal direction methods on the Tuebingen pairs are not comparable due to differing protocols (pair subsets, weightings, model selection, and abstention rates). It conducts a same-hands re-evaluation forcing binary decisions on the identical 102 pairs with no tuning, introduces a zero-parameter sorted-conditional compression baseline (quantized, sorted, first-differenced data fed to bz2), and reports that this baseline reaches 74.7% weighted accuracy (p=3.7e-7) while methods cluster in the low-to-mid 70s. It reproduces lower figures for RECI (70.7%) and forced-decision SLOPE (77.2%), attributes higher published numbers to test-set selection and significance-gated abstention, and releases code, pre-registrations, and per-pair outputs.

Significance. If the re-implementations and protocol hold, the work supplies a transparent, fully reproducible parameter-free baseline together with independent statistical tests (McNemar, p-values) and a confounding flag via compression scores (p=2.8e-68). The public release of code, pre-registrations, and per-pair outputs is a clear strength that enables direct verification and improves benchmarking standards in causal discovery.

major comments (1)
  1. [Abstract] Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler.
minor comments (1)
  1. [Abstract] Abstract: The brief description of the baseline ('quantized, sorted, first-differenced data') leaves the exact quantization scheme and differencing order implicit; a one-sentence expansion would improve standalone readability even though code is released.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the constructive comment regarding the abstract and the choice of protocol. We address the point directly below.

read point-by-point responses
  1. Referee: [Abstract] Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler.

    Authors: We agree that the forced binary-decision protocol is the foundation of our comparability claim and that it necessarily lowers SLOPE from its published decided-subset figure (which we already reproduce at 81.7 %). Our manuscript explicitly contrasts the two figures and attributes the difference to significance-gated abstention. Nevertheless, the referee's request for an explicit sensitivity analysis is reasonable. In the revised manuscript we will add a dedicated subsection that (i) scores all abstentions as errors (zero contribution to accuracy) for every method that permits them and (ii) applies proper scoring rules (Brier score and log-loss) to any probabilistic outputs that are available. This will allow readers to verify whether the low-to-mid-70s cluster persists under alternative treatments of abstention. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; baseline and evaluations are independent of fitted parameters or self-citation chains.

full rationale

The paper's central claims rest on a parameter-free compression baseline (sorted-conditional compression using bz2 on quantized, sorted, first-differenced data) applied to the external Tuebingen dataset under a fixed protocol of forced binary decisions on all 102 pairs. Accuracies (e.g., 74.7% weighted) and p-values (e.g., 3.7e-7) are computed directly from these runs and standard statistical tests (McNemar) without any parameter fitting, model selection on the test set, or reduction to quantities defined by the authors' prior work. No self-citations are load-bearing for the core result; the re-evaluations of other methods (SLOPE, RECI) reproduce published outputs or stored decisions on the same fixed pairs. The derivation chain is self-contained against external benchmarks and does not exhibit self-definitional, fitted-input, or ansatz-smuggling patterns.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The central claim rests on the validity of the Tuebingen pairs as representative cause-effect examples and on the assumption that compression of sorted first differences can serve as a causal direction signal. No free parameters or invented entities are introduced.

assumptions (2)
  • domain assumption The Tuebingen cause-effect pairs constitute a suitable fixed benchmark for comparing bivariate causal direction methods.
    All reported accuracies and comparisons depend on this dataset being accepted as ground truth.
  • ad hoc to paper Forcing a decision on every pair without abstention produces a fairer comparison than significance-gated protocols.
    This modeling choice defines the common ruler but is not derived from prior literature.

how reviews work

0 comments
Cite this review

Pith. "Pith review of One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline." pith.science (2026). https://pith.science/paper/DK6PCOQP

@misc{pith2026260623767,
  author       = {Pith},
  title        = {Pith review of: One Ruler: A Same-Hands Re-Evaluation of Bivariate Causal Direction on Tuebingen, with a Parameter-Free Compression Baseline},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DK6PCOQP}},
  note         = {Machine review of arXiv:2606.23767}
}
read the original abstract

Headline accuracies on the Tuebingen cause-effect pairs are routinely compared across papers even though each is measured under its authors' own protocol -- different pair subsets, weightings, model-selection, and decision rates. We argue this is the wrong comparison and run the right one: a same-hands re-evaluation in which every method is run by us on the identical 102 pairs, with one strict rule -- no tuning and a decision forced on every pair. As a clean reference point we introduce a deliberately minimal baseline: sorted-conditional compression, which feeds quantized, sorted, first-differenced data to an off-the-shelf compressor (bz2) and has zero fitted parameters. Under the common ruler the ranking differs sharply from the literature. Our baseline reaches 74.7% weighted accuracy (p = 3.7e-7); on the same 100 pairs that SLOPE is evaluated on it scores 76.0%, a 1.2-point gap below the authors' own forced-decision SLOPE (77.2%) that is well inside noise (McNemar p = 0.39). A faithful re-run of RECI lands at 70.7% -- inside the original authors' reported error bar, not the 77.5% often quoted (which we trace to a mis-copied cell). SLOPE's published 82.4% is a decided-subset figure: scoring the authors' own stored output only on the pairs its significance test chose to answer reproduces 81.7%. Under the common ruler the methods cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them. We document the mechanisms that inflate published figures (test-set model selection, significance-gated abstention) and contribute two further results: compression score magnitude is a model-free confounding flag (p = 2.8e-68), and a pre-registered falsification test fails in an instructive way that bounds the method's theoretical interpretation. Code, pre-registrations, and per-pair outputs are released.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

7 extracted references · 1 canonical work pages

  1. [1]

    Mooij, J

    J. Mooij, J. Peters, D. Janzing, J. Zscheischler, B. Schölkopf. Distinguishing cause from effect using observational data: methods and benchmarks.JMLR17(32):1–102, 2016

  2. [2]

    Janzing, B

    D. Janzing, B. Schölkopf. Causal inference using the algorith- mic Markov condition.IEEE Trans. Inf. Theory56(10):5168– 5194, 2010

  3. [3]

    Lemeire, D

    J. Lemeire, D. Janzing. Replacing causal faithfulness with algorithmic independence.Minds and Machines23(2):227– 249, 2013

  4. [4]

    Janzing et al

    D. Janzing et al. Information-geometric approach to inferring causal directions.Artificial Intelligence182–183:1–31, 2012

  5. [5]

    Blöbaum, D

    P. Blöbaum, D. Janzing, T. Washio, S. Shimizu, B. Schölkopf. Cause-effect inference by comparing regression errors.AIS- TATS2018; extended analysis,PeerJ CS2019. 5

  6. [6]

    A. Marx, J. Vreeken. Telling cause from effect using MDL- based local and global regression.ICDM2017

  7. [7]

    Hlavackova-Schindler, A

    K. Hlavackova-Schindler, A. Marsela. Identifying causal di- rection via dense functional classes. arXiv:2509.00538, 2025. 6

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.