{"id":"ad6a2ad1-d45b-46f6-8584-4262a08b5003","arxiv_id":"2606.23767","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Uniform forced-decision re-evaluation on 102 Tuebingen pairs shows a parameter-free sorted-conditional compression baseline at 74.7% weighted accuracy, tying methods in the low-to-mid 70s.","lead":"The paper re-evaluates causal direction methods on the Tuebingen dataset using identical 102 pairs, forced decisions, and no tuning for every method. It introduces a zero-parameter compression baseline that matches top published methods under this uniform protocol.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Forcing binary decisions on every pair may not be the appropriate common ruler, as it can disadvantage methods designed around significance-based abstention.","rationale":"The reader's weakest_assumption exactly matches the load-bearing premise identified above. No stronger internal inconsistency (e.g., in the re-implementation details, p-value calculations, or baseline construction) was located that would independently threaten the central numerical claim once the protocol premise is granted.","tokens_in":1925,"tokens_out":351,"duration_ms":24016,"concrete_test":"For each method with an internal significance test, re-score both the method and the baseline only on the subset of pairs the method chooses to answer; if the method's accuracy on its decided subset exceeds the baseline's accuracy on the same subset by more than the McNemar noise level reported in the paper, the forced-decision equalization weakens.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim (baseline at 74.7% weighted accuracy ties the strongest methods under the common ruler) requires that the strict forced-decision protocol on the fixed 102 pairs is the correct basis for comparison. This protocol is least secure for methods (e.g., SLOPE) whose original design and reported figures rely on internal significance tests to abstain; the paper itself shows SLOPE's published 82.4% is a decided-subset figure (reproduced at 81.7%), while the forced-decision re-run drops to 77.2%. If abstention is a legitimate part of the method, the uniform forced-decision ruler systematically alters relative performance and does not demonstrate that the methods are no stronger than the compressor once their native decision rules are respected.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript argues that published accuracies for bivariate causal direction methods on the Tuebingen pairs are not comparable due to differing protocols (pair subsets, weightings, model selection, and abstention rates). It conducts a same-hands re-evaluation forcing binary decisions on the identical 102 pairs with no tuning, introduces a zero-parameter sorted-conditional compression baseline (quantized, sorted, first-differenced data fed to bz2), and reports that this baseline reaches 74.7% weighted accuracy (p=3.7e-7) while methods cluster in the low-to-mid 70s. It reproduces lower figures for RECI (70.7%) and forced-decision SLOPE (77.2%), attributes higher published numbers to test-set selection and significance-gated abstention, and releases code, pre-registrations, and per-pair outputs.","tokens_in":2112,"tokens_out":490,"duration_ms":26268,"significance":"If the re-implementations and protocol hold, the work supplies a transparent, fully reproducible parameter-free baseline together with independent statistical tests (McNemar, p-values) and a confounding flag via compression scores (p=2.8e-68). The public release of code, pre-registrations, and per-pair outputs is a clear strength that enables direct verification and improves benchmarking standards in causal discovery.","major_comments":[{"comment":"Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler.","section":"Abstract"}],"minor_comments":[{"comment":"Abstract: The brief description of the baseline ('quantized, sorted, first-differenced data') leaves the exact quantization scheme and differencing order implicit; a one-sentence expansion would improve standalone readability even though code is released.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comment regarding the abstract and the choice of protocol. We address the point directly below.","responses":[{"response":"We agree that the forced binary-decision protocol is the foundation of our comparability claim and that it necessarily lowers SLOPE from its published decided-subset figure (which we already reproduce at 81.7 %). Our manuscript explicitly contrasts the two figures and attributes the difference to significance-gated abstention. Nevertheless, the referee's request for an explicit sensitivity analysis is reasonable. In the revised manuscript we will add a dedicated subsection that (i) scores all abstentions as errors (zero contribution to accuracy) for every method that permits them and (ii) applies proper scoring rules (Brier score and log-loss) to any probabilistic outputs that are available. This will allow readers to verify whether the low-to-mid-70s cluster persists under alternative treatments of abstention.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The claim that the compressor 'ties the strongest of them' under the common ruler rests on the forced binary-decision protocol applied to all 102 pairs. This protocol changes SLOPE performance from its published decided-subset figure (82.4%, reproduced at 81.7%) to 77.2% forced-decision; a sensitivity check that scores methods allowing abstention (with undecided pairs scored as errors or via proper scoring rules) is needed to confirm that the low-to-mid-70s clustering is not an artifact of the chosen ruler."}],"tokens_in":1619,"tokens_out":336,"duration_ms":18185,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main point is that published accuracies on the Tuebingen pairs cannot be compared directly because each paper used its own subset, weighting, tuning, or abstention rule. They fix that by re-running everything themselves on the identical 102 pairs with a single rule: no tuning and a decision on every pair. As reference they add a deliberately minimal baseline that quantizes, sorts, differences, and compresses with bz2, zero fitted parameters.\n\nThis is the useful part. They show the re-run numbers: their compressor at 74.7% weighted accuracy, SLOPE forced-decision at 77.2% on the overlapping pairs (inside noise of the compressor), and RECI at 70.7% once the mis-copied cell is fixed. They also document that SLOPE's 82.4% figure only scores the pairs its significance test answered. Code, pre-registrations, and per-pair outputs are released, so the claims can be checked.\n\nThe soft spot is the protocol choice itself. Forcing a binary answer on every pair is clean for comparison, but it changes the task for methods built around abstention. The paper treats the forced version as the correct common ruler; whether that matches how people actually want to use the methods is left to the reader. The additional result on compression magnitude as a confounding flag is straightforward but secondary.\n\nThis is for researchers who cite or run bivariate causal direction benchmarks and want to know how much of the reported performance depends on evaluation choices. The concrete numbers, statistical tests, and released artifacts make it worth a serious referee's time to verify the re-implementations.\n\nSend it to peer review. The re-evaluation is checkable and addresses a real inconsistency in the literature.","headline":"Under one forced-decision protocol on all 102 pairs the methods land in the low-to-mid 70s and a zero-parameter compressor matches the strongest of them.","tokens_in":2580,"tokens_out":437,"would_cite":false,"duration_ms":18310,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A zero-parameter compression baseline reaches 74.7 percent weighted accuracy on all 102 Tuebingen pairs when every method must decide without tuning or abstention.","keywords":["causal discovery","Tuebingen cause-effect pairs","bivariate causal direction","compression baseline","re-evaluation protocol","parameter-free method","forced decision evaluation"],"falsifier":"A re-run of the same 102 pairs under the forced-decision protocol in which any literature method exceeds the baseline by a margin larger than the McNemar test noise level would falsify the reported clustering and tie.","tokens_in":2807,"feed_emoji":"📊","tokens_out":880,"duration_ms":25877,"temperature":0.7,"pith_summary":"The paper establishes that routine comparisons of bivariate causal direction methods on the Tuebingen cause-effect pairs rest on inconsistent protocols that differ in pair subsets, weightings, model selection, and decision rates. It therefore applies one uniform protocol: every method is re-run on the identical 102 pairs with no tuning permitted and a binary decision required for every pair. Under this common ruler a deliberately minimal baseline that sorts, quantizes, first-differences the data and feeds it to an off-the-shelf bz2 compressor scores 74.7 percent weighted accuracy. Re-evaluations of published methods show that their higher headline numbers often arise from scoring only on decided subsets or from test-set model selection. The result is that accuracies cluster tightly in the low-to-mid 70s, with the parameter-free compressor tying the strongest competitors.","feed_headline":"Zero-parameter compressor ties top methods at 74.7% on Tuebingen pairs","feed_subtitle":"When every algorithm must answer all 102 pairs without tuning or skipping, published figures converge and a simple baseline matches the lead","key_machinery":"The same-hands re-evaluation protocol that applies one strict rule—no tuning and a decision forced on every pair—to all methods on the full set of 102 Tuebingen pairs, benchmarked against a sorted-conditional compression baseline that feeds quantized, sorted, first-differenced data to an off-the-shelf bz2 compressor.","core_discovery":"Under the common ruler of evaluating every method on the identical 102 pairs with forced decisions and no tuning, the sorted-conditional compression baseline reaches 74.7 percent weighted accuracy. A faithful re-run of RECI lands at 70.7 percent. SLOPE's published 82.4 percent is reproduced only when scoring is restricted to the pairs its significance test chooses to answer; on the full set the figure drops. The methods therefore cluster in the low-to-mid 70s and the zero-parameter compressor ties the strongest of them.","pith_inferences":["Adopting the forced-decision protocol more widely would require future causal discovery papers to report both full-set and selective accuracies for direct comparison.","The observed clustering implies that further gains from increasingly complex models may be small once protocol differences are eliminated, shifting attention to data characteristics.","The same-hands approach could be applied to other bivariate or low-dimensional causal benchmarks to test whether the Tuebingen convergence generalizes.","A compression baseline of this form supplies an immediate, training-free reference point for any new method proposed on similar pair data."],"forward_implications":["SLOPE's published 82.4 percent reflects performance only on the subset its significance test elects to answer rather than on the full set.","RECI's re-run score of 70.7 percent falls inside the original authors' reported error bar, not the 77.5 percent figure often quoted.","Compression score magnitude functions as a model-free indicator of confounding (p = 2.8e-68).","A pre-registered falsification test fails in a manner that bounds the theoretical interpretation of the compression approach.","Under the uniform protocol all examined methods perform in the low-to-mid 70s."],"fun_headline_variants":["Same-hands Tuebingen run: compressor hits 74.7% on all 102 pairs","Zero-param baseline ties top at 74.7% under identical forced rules","Full 102-pair re-evaluation puts methods in 70-75% range","RECI faithful rerun scores 70.7% on strict Tuebingen protocol","SLOPE falls to 77.2% on full set from published 82.4%"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The claim rests on the premise that forcing a binary decision on every pair without allowing significance-based abstention or model selection constitutes the correct common ruler for comparison.","fun_headline_variants_meta":{"raw":{"variants":["Same-hands Tuebingen run: compressor hits 74.7% on all 102 pairs","Zero-param baseline ties top at 74.7% under identical forced rules","Full 102-pair re-evaluation puts methods in 70-75% range","RECI faithful rerun scores 70.7% on strict Tuebingen protocol","SLOPE falls to 77.2% on full set from published 82.4%"]},"model":"grok-4.3","cost_usd":0.002567,"raw_usage":{"total_tokens":1580,"prompt_tokens":891,"num_sources_used":0,"completion_tokens":109,"cost_in_usd_ticks":25674500,"prompt_tokens_details":{"text_tokens":891,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":580,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":891,"tokens_out":109,"duration_ms":4977,"temperature":1.0,"reasoning_tokens":580,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T09:14:50.055331+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A re-run of the same 102 pairs under the forced-decision protocol in which any literature method exceeds the baseline by a margin larger than the McNemar test noise level would falsify the reported clustering and tie.","supporting_citations":[],"review_version":1}