{"total":10,"items":[{"citing_arxiv_id":"2607.01104","ref_index":89,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"CausalMix: Data Mixture as Causal Inference for Language Model Training","primary_cat":"cs.LG","submitted_at":"2026-07-01T15:56:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"CausalMix fits a causal model on 512 runs of a 0.5B model to estimate CATE, then extrapolates optimal mixtures for an 800K data pool applied to 7B and 4B models, outperforming RegMix.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.22335","ref_index":122,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Learning Causal Orderings for In-Context Tabular Prediction","primary_cat":"cs.LG","submitted_at":"2026-05-21T11:22:56+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"TabOrder learns unsupervised causal variable orderings and enforces them with order-constrained attention for tabular prediction and imputation under distribution shifts.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.19343","ref_index":42,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"What Makes a Representation Good for Single-Cell Perturbation Prediction?","primary_cat":"cs.LG","submitted_at":"2026-05-19T04:30:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"PerturbedVAE disentangles perturbation-specific signals from invariant gene expression structure to recover causal representations and improve out-of-distribution prediction in single-cell perturbation modeling.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"(i)Invertibility and Smoothness:The unknown nonlinear mappinggis smooth and invertible. (ii)Environmental Sufficiency:There exist2d ν distinct environments{u 1, . . . ,um}relative to a referenceu 0 such that the matrix L⊤ = [∆η(u1), . . . ,∆η(u 2dν)]⊤ ∈R 2dν ×2dν (41) has full column rank2d ν, where (elementwise divisions) ∆η(u) :=   µν (u) βν (u) − µν (u0) βν (u0) − 1 2 \u0010 1 βν (u) − 1 βν (u0) \u0011   ∈R 2dν .(42) . (iii)Optimal Alignment:The alignment loss, e.g., Eq. 4 attains its global minimum such thatf ι(x(u)) =f ι(x(u0))almost surely for anyu,u 0, wheref ι = ˆ g−1 ι . (iv)Intervention Sufficiency:The function class ofλsatisfies the following condition: there existsu i, such that, for all parent nodesz j ∈pa i ofz i,λ j,i = 0. Then, the true latent causal variableszare related to the variables ˆzestimated by matching likelihood (i."},{"citing_arxiv_id":"2605.19313","ref_index":16,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"A Unified Framework for Structure-Aware Clustering and Heterogeneous Causal Graph Learning","primary_cat":"stat.ML","submitted_at":"2026-05-19T03:48:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DAG-DC-ADMM jointly clusters subjects and learns their cluster-specific causal DAGs via structural equation modeling, groupwise truncated Lasso fusion penalties, and an ADMM solver for the resulting nonconvex problem.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18633","ref_index":18,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Stable Causal Discovery via Directed Acyclic Graph Aggregation","primary_cat":"stat.ME","submitted_at":"2026-05-18T16:41:31+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DAGgr aggregates weighted candidate DAGs using out-of-sample predictive likelihood and an acyclicity-preserving threshold, with claimed finite-sample bounds and consistency, outperforming baselines in simulations and protein network data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.17465","ref_index":13,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"TriOpt: A Scalable Algorithm for Linear Causal Discovery","primary_cat":"cs.LG","submitted_at":"2026-05-17T14:07:27+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"TriOpt recovers topological order via Sherman-Morrison downdates on linear kernels then solves a convex program for the DAG edges, claiming exact recovery under the true order and large speedups on high-dimensional data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.06315","ref_index":16,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"End-to-End Identifiable and Consistent Recurrent Switching Dynamical Systems","primary_cat":"stat.ML","submitted_at":"2026-05-07T14:14:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Identifiability is proven for recurrent nonlinear switching dynamical systems under flexible assumptions, and ΩSDS is introduced as a flow-based estimator that improves disentanglement and forecasting over VAE-based methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.04381","ref_index":58,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Causal discovery under mean independence and linearity","primary_cat":"stat.ME","submitted_at":"2026-05-06T01:16:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LiMIAM and DirectLiMIAM enable causal discovery from observational data under mean-independent but dependent disturbances, outperforming LiNGAM in simulations and recovering plausible orderings in oil market data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.03178","ref_index":20,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"Structure Learning for Directed Trees with Zero-Inflated Compositional Nodes","primary_cat":"stat.ME","submitted_at":"2026-05-04T21:41:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A new directed tree structure learning framework for zero-inflated compositional nodes uses KL divergence scoring and column-stochastic transition matrices for conditional expectations, with proven consistency and finite-sample guarantees.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.01134","ref_index":135,"ref_count":1,"confidence":0.35,"is_internal_anchor":false,"paper_title":"To Use AI as Dice of Possibilities with Timing Computation","primary_cat":"cs.AI","submitted_at":"2026-05-01T22:25:29+00:00","verdict":"REJECT","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"The paper defines possibility space, timing computation, and causal factum to make timing a computable variable, and illustrates the framework with automatic trajectory discovery and counterfactual timing on 3,276 breast-cancer patients.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}