{"id":"9e1561e0-b060-4ed7-8e22-16ec6e30e36f","arxiv_id":"2607.14124","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using DirectLiNGAM to pick 'causal' school features before hierarchical clustering yields clusters that are more compact and more strongly linked to Saeb performance than clusters from all 33 variables.","lead":"This paper tests whether choosing school features through causal discovery before clustering produces more meaningful school groupings than using all available features. It finds the causal-feature clusters are better separated and more strongly tied to test scores, but the comparison lacks a critical control.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal scenario is compared only to the full 33-variable baseline; the Markov-blanket subset's better silhouette/η² may be a dimensionality artifact. A same-size non-causal baseline is needed before any causal advantage is claimed.","rationale":"The strongest claim is that causal feature selection, via the Markov blanket, improves typology quality and performance discrimination. For that claim to hold, the improvement must be due to the causal provenance of the selected variables, not merely to using fewer, more outcome-related variables. The paper's baseline is always the full 33-variable input; no same-size comparison is reported. This is the least secure condition because it is necessary regardless of causal discovery correctness: even a perfectly estimated DAG whose Markov blanket yields m variables would produce the same experimental inequality. The proposed random/correlational baseline test directly addresses this. I agree with the reader that the post-hoc choice among four causal algorithms and the unavailability of code/data are concerns, but those are secondary: the subset-size confound alone undermines the comparative causal claim. I do not think the paper should be rejected: the observed effect sizes are plausible, the methods are standard, and the authors acknowledge limitations in Section 7. But the conclusion that causal reasoning 'demonstrates' an enhancement (Section 7) overstates what the design supports; the paper should be CONDITIONAL on adding a non-causal same-size control or weakening the causal interpretation. Since the reader's verdict is already CONDITIONAL, no verdict change is needed.","tokens_in":9740,"tokens_out":4493,"duration_ms":47340,"concrete_test":"Re-run the clustering pipeline of §4.3 with k=8, Ward linkage, and Euclidean distance on (a) 1000 random subsets of the 33 variables of size m = |MB| and (b) the m variables with the largest absolute Pearson/Spearman correlation with the weighted Saeb score. Compare silhouette score, Davies–Bouldin index, and ANOVA η² to the causal scenario's values (0.3496, 1.2040, 0.133). If the causal values are not above the 95th percentile of distributions (a) and (b), the reported advantage is a dimensionality/subset-size effect, not evidence for causal feature selection.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing comparison in the paper is between two unequal feature sets: the baseline clustering uses all 33 variables (§4.3), while the causal scenario uses only the Markov blanket of the weighted Saeb score, extracted from a DirectLiNGAM DAG that was chosen after the fact as 'the most coherent structure' (§5.1). The reported improvements—silhouette 0.1211→0.3496 and Davies–Bouldin 2.0120→1.2040 in Table 4, η² 0.085→0.133 and Kruskal–Wallis 67.48→109.76 in Table 3—are therefore not attributable to causality without a control for subset size and dimensionality. In hierarchical clustering with Euclidean distance and Ward's linkage, adding many weakly relevant or noisy coordinates flattens distance ratios and can worsen all geometric indices; restricting to a smaller set, especially one selected using the outcome itself, would be expected to improve cluster separation and performance association even under a purely correlational feature-selection rule. The Markov blanket is by definition the optimal predictive set for the target (§3.2), so its better outcome association is a property of predictive feature selection, not of causal structure. Thus the central claim 'clustering based on causal features improves the discriminative capacity' (§5.3) is confounded: the design cannot distinguish 'causal' from 'fewer, outcome-correlated features.' This concern is independent of whether the DirectLiNGAM graph is correct; it would invalidate the causal interpretation even if the graph were perfect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares two hierarchical clustering pipelines for Brazilian public schools: a baseline using all 33 census-derived variables and a 'causal' scenario using only the Markov blanket of a weighted Saeb performance score, where the blanket is extracted from a DirectLiNGAM causal DAG. The authors report that the causal scenario yields higher silhouette score (0.1211 to 0.3496), lower Davies–Bouldin index (2.0120 to 1.2040), larger ANOVA F (10.29 to 16.97), larger eta-squared (0.085 to 0.133), and larger Kruskal–Wallis statistic (67.48 to 109.76), concluding that causal feature selection improves the discriminative capacity, performance association, and internal consistency of school typologies. The manuscript also includes structural similarity metrics (ARI, AMI, Cramér's V) between the two partitions.","tokens_in":10039,"tokens_out":1981,"duration_ms":20490,"significance":"If the central claim were established, the paper would contribute a practical template for integrating causal discovery with clustering in educational data, with potential policy relevance for identifying school typologies that align with performance. Strengths include the use of multiple complementary validation metrics, clearly reported effect sizes and test statistics, and a transparent comparison framework. However, the current design does not support the causal interpretation because the improvement could arise from reducing the feature set size or from selecting features by their association with the outcome, rather than from causal structure per se. The paper would be strengthened by including a same-size non-causal baseline, external validation, and robustness checks across causal discovery algorithms. These are feasible fixes within the manuscript's scope, so the result is defensible only after substantial revision.","major_comments":[{"comment":"The central comparison is confounded by feature-set size. The baseline uses all 33 variables, while the causal scenario uses only the Markov blanket of the Saeb score, which is likely a smaller subset. In hierarchical clustering with Euclidean distance and Ward linkage, adding many weakly relevant or noisy coordinates can flatten distance ratios and degrade geometric indices; restricting to a smaller feature set can improve silhouette and Davies–Bouldin even when the features are chosen by simple correlation with the target. To support the claim that causality—rather than dimensionality—drives the improvement, the authors must report a control clustering using the same number of features selected at random or by a purely associational criterion (e.g., top-k by absolute correlation with Saeb). Without this control, the improvements in Tables 3 and 4 are not attributable to causal structur","section":"§4.3, Tables 3 and 4"},{"comment":"The Markov blanket is described as 'the theoretically optimal feature set for predicting T' and is extracted using the same Saeb outcome variable used to evaluate cluster-performance association. This creates a selection loop: features are chosen for their statistical relation to the outcome, and then cluster-outcome association is measured on the same data. The better ANOVA/Kruskal–Wallis and effect sizes in the causal scenario are expected from predictive feature selection, not from causal structure. An external or held-out evaluation—e.g., estimating the Markov blanket on one subsample and evaluating clustering on another—would mitigate this circularity. The current design cannot distinguish 'causal' from 'outcome-correlated, smaller feature set.'","section":"§4.2, §3.2, Eq. (1)"},{"comment":"The causal DAG is selected because it 'exhibited the most coherent structure' among DAGMA, FCI, GES, and DirectLiNGAM, but no formal criterion or sensitivity analysis is provided. This after-the-fact selection is vulnerable to confirmation bias and does not quantify the uncertainty in the recovered graph. The authors should report the Markov blankets obtained from the other algorithms, the stability of the subsequent clustering results, and the reason DirectLiNGAM's output is preferred (e.g., agreement with established educational findings, stability across subsamples). Without this, the 'causal' scenario is effectively one arbitrary feature subset among many possible ones.","section":"§5.1, §4.2"},{"comment":"DirectLiNGAM relies on linearity, non-Gaussian disturbances, acyclicity, and causal sufficiency (no latent confounders). The paper does not test or discuss these assumptions for the educational dataset, where latent constructs (e.g., school management quality, community engagement) are plausible. If any assumption fails, the estimated DAG—and hence the Markov blanket—may not be causally valid, making the 'causal' scenario an associational feature subset. At minimum, the authors should acknowledge this threat and provide diagnostic checks (e.g., testing pairwise linearity, comparing LiNGAM with a non-linear method) or a robustness analysis.","section":"§3.1, §5.1"}],"minor_comments":[{"comment":"The section title 'Casual Features' appears to be a typo for 'Causal Features'.","section":"§4.2 title"},{"comment":"The table uses Portuguese dimension names (e.g., 'Infraestrutura Básica', 'Recursos e Espaços'). English translations or a bilingual table would improve accessibility.","section":"Table 1"},{"comment":"References [25] and [26] appear to duplicate the same paper (Weinstein and Blei, 'Hierarchical Causal Models'). Please consolidate or clarify.","section":"References"},{"comment":"The sentence 'the inclusion of causal features does not merely generate an alternative partition' is imprecise; the comparison is between feature sets, not 'inclusion' of features. Rewording would avoid confusion.","section":"§5.5"},{"comment":"The School Census is described as not public and used under institutional access, but no data availability statement or replication details (e.g., software versions, hyperparameters for DirectLiNGAM) are provided. Adding these would strengthen reproducibility.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising empirical setup and clear reporting, but the central causal claim is currently underidentified. The missing same-size non-causal control is the decisive issue; if the authors add it and it shows no advantage for causal features, the manuscript's main conclusion collapses. If the control is added and the causal advantage persists, the paper could become a solid applied contribution. The subjectivity in graph selection and lack of external validation are also concerns, but they can be addressed within a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Main thing to know: this is an applied paper that runs a reasonable pipeline — DirectLiNGAM, Markov blanket extraction, Ward hierarchical clustering — on Brazilian School Census and Saeb data, and finds better cluster metrics for the Markov-blanket variable set than for all 33 variables. Taken at face value, it's a modest, useful demonstration that feature selection can sharpen clusters. But the paper's own claim is stronger: that the improvement is due to causal features. The design does not support that.\n\nThe comparison is 33 variables versus a small Markov-blanket set, with no control for subset size. Any small, outcome-correlated feature set would likely improve silhouette and ANOVA separation, so the 'causal' label is doing the interpretive work, not the data. On top of that, the causal graph was chosen after the fact as the 'most coherent' among four algorithms, which adds subjectivity. The conclusion says the approach 'demonstrates' an enhancement — that overstates the evidence.\n\nCredit where earned: the authors know the methods, the evaluation framework is reasonably comprehensive (ANOVA, Kruskal–Wallis, silhouette, Davies–Bouldin, ARI/AMI), and the limitations section is honest about single-dataset generalizability. The specific empirical comparison on 2023 census plus Saeb data appears to be new, and as an application it is legitimate. The paper is clearly written and the clustering analyses are easy to follow.\n\nSoft spots, in order of severity:\n(1) The central causal claim is confounded; a same-size baseline using random or associational features is needed before any causal advantage can be claimed.\n(2) The graph selection is post hoc; no stability analysis or sensitivity to algorithm choice is reported.\n(3) DirectLiNGAM assumptions — linearity, non-Gaussianity, no latent confounders — are asserted rather than checked, and location was deliberately excluded even though it could be a confounder.\n(4) The Markov blanket variables are not listed, so readers can't see what actually changed.\n(5) Data and code are not available, which limits replication.\n\nThis paper is for applied educational data mining researchers who want a template for combining causal discovery with clustering. It deserves a serious referee because the empirical finding is plausible and could be interesting if the control issue were fixed. I'd send it to peer review, but the review should require a same-size non-causal baseline and ideally a holdout evaluation of cluster-performance association. If the causal advantage survives that, it becomes a useful contribution.","headline":"A plausible application with an unsupported causal claim: the clustering improvement may just be a smaller, outcome-correlated feature set, and the causal graph was chosen post hoc.","tokens_in":10577,"tokens_out":2240,"would_cite":false,"duration_ms":23360,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62H30","62P25"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that clustering schools on causally relevant features, rather than all available census variables, produces typologies that separate educational performance more sharply and are geometrically more coherent.","keywords":["educational data mining","hierarchical clustering","causal discovery","Markov blanket","DirectLiNGAM","school performance","Saeb","Brazilian schools"],"falsifier":"Re-run the clustering pipeline with random subsets of the same size as the causal feature set, many times. If random subsets frequently match or beat the causal scenario's silhouette (about 0.35) and ANOVA effect size (eta-squared about 0.13), then the causal interpretation is not needed to explain the improvement.","tokens_in":103,"feed_emoji":"🎓","tokens_out":6746,"duration_ms":117008,"temperature":0.7,"pith_summary":"This paper tries to show that injecting causal information into clustering produces school typologies that are more useful than those built from raw descriptive variables. The authors cluster Brazilian public schools with hierarchical clustering under two scenarios: one using all 33 school-census variables, and one using only the variables that a causal model places in the Markov blanket of the school's Saeb performance score. They find that the causal clusters separate high- and low-performing schools more sharply (ANOVA F 10.29 to 16.97; eta-squared 0.085 to 0.133), are more internally compact and separated (silhouette 0.12 to 0.35; Davies-Bouldin 2.01 to 1.20), and satisfy variance homogeneity where the baseline does not. The conclusion is that causal feature selection, not just statistical association, makes school typologies more interpretable and more relevant for policy analysis.","feed_headline":"Causal features make school typologies separate test scores better","feed_subtitle":"Clustering Brazil's public schools on the causally relevant subset of 33 census variables sharply improves separation of Saeb performance.","key_machinery":"The Markov blanket of the target variable: within a fitted causal directed acyclic graph, the set of a variable's parents, children, and co-parents, which is the minimal set that screens the target off from all other variables. The paper uses the Markov blanket as a causally grounded feature selector: after fitting a linear non-Gaussian acyclic causal model (DirectLiNGAM), it retains only the blanket variables for clustering. The clustering itself is hierarchical agglomerative clustering with Euclidean distance and minimum-variance linkage, producing eight clusters from the selected feature space.","core_discovery":"The paper's central claim is that clustering schools on the Markov blanket of the Saeb score, rather than on all 33 census variables, produces school typologies that are more compact, better separated, and more strongly associated with measured performance. Using a causal graph estimated by DirectLiNGAM and extracting the Markov blanket of the performance target, the authors built an eight-cluster hierarchical solution and compared it with a baseline using all variables. The causal solution wins on every quantitative criterion the paper reports: larger ANOVA F and non-parametric test statistics, larger effect sizes, homogeneity of variances, better silhouette score, and lower Davies-Bouldin","pith_inferences":["Inference: the paper does not run an ablation against random subsets of the same size; if random subsets reproduced the silhouette and effect-size gains, the advantage would be attributable to dimension reduction rather than to causal validity.","Inference: the authors selected the 'most coherent' DAG among four causal-discovery algorithms after inspecting the results, so the causal feature set is partly a matter of judgment; checking Markov blankets from the other algorithms would show how much the typology improvement depends on that choice.","Inference: the weighted Saeb score blends elementary and high school stages, and if causal mechanisms differ by stage, stage-specific blankets might yield even sharper typologies and different policy levers."],"forward_implications":["If the causal scenario is correct, schools grouped by causal features are more homogeneous in performance and more distinct across groups, making cluster membership a stronger signal for diagnosing school effectiveness.","The causal clusters satisfy the homogeneity-of-variances assumption that the baseline violates, so comparisons of mean performance across clusters are statistically more trustworthy.","Moderate overlap between the two partitions (ARI 0.338, AMI 0.470) implies that causal selection reorganizes schools rather than reproducing existing typologies, so policy conclusions would differ between the two schemes.","Because the causal feature space yields better geometric quality, the typologies are better suited as targets for interventions tailored to distinct school profiles.","The authors' stated next step, replicating the analysis nationwide, would test whether the causal blanket found here generalizes beyond their dataset."],"fun_headline_variants":["Causal subset beats all 33 variables for school clusters","Markov blanket clustering sharpens school test score separation","Clustering on causal features yields better school typologies","Causal variable selection improves cluster separation in schools","School clusters separated better using causal Markov blanket"],"cache_read_input_tokens":11776,"weakest_assumption_plain":"The claim stands or falls on whether the causal graph the authors chose really represents how school characteristics cause performance; if that graph is wrong, the causal features are just another arbitrary subset and the observed advantages could be an artifact of using fewer variables.","fun_headline_variants_meta":{"raw":{"variants":["Causal subset beats all 33 variables for school clusters","Markov blanket clustering sharpens school test score separation","Clustering on causal features yields better school typologies","Causal variable selection improves cluster separation in schools","School clusters separated better using causal Markov blanket"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000126,"raw_usage":{"total_tokens":914,"prompt_tokens":677,"completion_tokens":237,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":421,"completion_tokens_details":{"reasoning_tokens":163}},"tokens_in":421,"tokens_out":237,"duration_ms":3287,"temperature":1.0,"reasoning_tokens":163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T10:40:51.550381+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the clustering pipeline with random subsets of the same size as the causal feature set, many times. If random subsets frequently match or beat the causal scenario's silhouette (about 0.35) and ANOVA effect size (eta-squared about 0.13), then the causal interpretation is not needed to explain the improvement.","supporting_citations":[],"review_version":1}