PROMETHEUS builds causal atlases from text and data using local predictive-state models and sheaf gluing to create navigable Topos World Models that expose evidence strength and coherence gaps.
hub
Causal reasoning and large language models: Opening a new frontier for causality
23 Pith papers cite this work, alongside 93 external citations. Polarity classification is still indexing.
hub tools
representative citing papers
TCD-Arena is a new customizable testing framework that runs millions of experiments to map how 33 different assumption violations affect time series causal discovery methods and shows ensembles can boost overall robustness.
PRCD-MAP assigns per-edge trust to imperfect priors in causal discovery via empirical Bayes calibration and MLP propagation, delivering an ε-safety guarantee that vanishes at prior-quality extremes and empirical gains on CausalTime datasets.
Proposes a sequential causal discovery framework integrating noisy LM priors with batch data via PAG representation and adaptive edge querying for improved structural accuracy.
MALLM-GAN uses multi-agent LLMs to emulate GAN architecture for generating higher-quality synthetic tabular data from small samples than prior models, while preserving privacy.
LLMs suppress causal caution in practical advisory contexts (rates drop from 91.7-100% to 6.7-18.3%) but recover it with a self-correction prompt (to 71.4-100%).
A world model is a positive semidefinite coupling kernel over admissible possible worlds, with the off-diagonal supplying the structural information for counterfactual queries that standard prediction cannot recover.
Lexical anonymization via Caliper causes consistent accuracy drops of 7-30 percentage points across LLMs on causal benchmarks, indicating reliance on lexical anchors rather than structural causal reasoning.
ORCA is an agent-orchestrated interactive copilot that automates and guides end-to-end causal analysis from workflow selection to report generation across real-world use cases.
CausalGuard aggregates LLM-proposed and data-pruned DAGs to weight doubly robust pseudo-outcomes and applies conformal calibration to deliver finite-sample marginal coverage for conditional average treatment effects under graph uncertainty.
Using FOMC minutes to propose regime-shift candidates and a lenient text check to ratify data-detected candidates, the pipeline reaches F1=0.82 on 26 monetary-policy anchors, beating every data-only baseline.
CIVeX maps agent tool calls to structural causal queries, checks identifiability, and issues auditable verdicts to prevent false executions while preserving utility on confounded benchmarks.
Four axioms (Causality, Minimality, Separability, Stability) are formalized for latent thought representations; audits of open LLMs on 23 tasks show none satisfy all four and representations add little beyond input embeddings.
Chain-of-thought helps LLMs on obvious policy findings but its benefit collapses on counter-intuitive ones, with case intuitiveness dominating model and prompt choice.
Introduces the CAUSALT3 benchmark for causal reasoning across Pearl's ladder and Regulated Causal Anchoring (RCA) to reduce sycophancy and skepticism in LLMs via inference-time verification.
Introduces CounterBench benchmark and CoIn iterative reasoning method showing LLMs perform near random on formal counterfactual tasks but improve substantially with guided backtracking.
CausalSynth combines structural causal models with LLMs and iterative verification to produce synthetic data that respects given causal structures while remaining linguistically natural.
Hume's causal judgment requires experiential grounding, structured retrieval, and vivacity transfer, conditions that Bayesian formalizations abstract away while LLMs retain only statistical updating.
Supervised learning across AI systems vindicates a uniform error-driven associationism for cognition, though operating inside advanced computational structures beyond classical associationist models.
The authors introduce a validation framework showing LLMs can pull causal links from disaster social media but require checks against post-event evidence to avoid relying on model priors.
Gemma 3 introduces multimodal open models with architectural changes for efficient long context, trained via distillation and a new post-training recipe that makes the 4B version competitive with prior 27B models and the 27B version comparable to Gemini-1.5-Pro.
ODYSSEY is a sheaf-theoretic framework for building verifiable foundation models as compositions of foundries via left and right Kan extensions.
citing papers explorer
-
PROMETHEUS: Automating Deep Causal Research Integrating Text, Data and Models
PROMETHEUS builds causal atlases from text and data using local predictive-state models and sheaf gluing to create navigable Topos World Models that expose evidence strength and coherence gaps.
-
TCD-Arena: Assessing Robustness of Time Series Causal Discovery Methods Against Assumption Violations
TCD-Arena is a new customizable testing framework that runs millions of experiments to map how 33 different assumption violations affect time series causal discovery methods and shows ensembles can boost overall robustness.
-
PRCD-MAP: Learning How Much to Trust Imperfect Priors in Causal Discovery
PRCD-MAP assigns per-edge trust to imperfect priors in causal discovery via empirical Bayes calibration and MLP propagation, delivering an ε-safety guarantee that vanishes at prior-quality extremes and empirical gains on CausalTime datasets.
-
Sequential Causal Discovery with Noisy Language Model Priors
Proposes a sequential causal discovery framework integrating noisy LM priors with batch data via PAG representation and adaptive edge querying for improved structural accuracy.
-
MALLM-GAN: Multi-Agent Large Language Model as Generative Adversarial Network for Synthesizing Tabular Data
MALLM-GAN uses multi-agent LLMs to emulate GAN architecture for generating higher-quality synthetic tabular data from small samples than prior models, while preserving privacy.
-
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
LLMs suppress causal caution in practical advisory contexts (rates drop from 91.7-100% to 6.7-18.3%) but recover it with a self-correction prompt (to 71.4-100%).
-
WorldKernel: A World Model is the Coupling Kernel of Admissible Possible Worlds
A world model is a positive semidefinite coupling kernel over admissible possible worlds, with the off-diagonal supplying the structural information for counterfactual queries that standard prediction cannot recover.
-
Caliper: Probing Lexical Anchors versus Causal Structure in LLMs
Lexical anonymization via Caliper causes consistent accuracy drops of 7-30 percentage points across LLMs on causal benchmarks, indicating reliance on lexical anchors rather than structural causal reasoning.
-
ORCA: An End-to-End Interactive Copilot for Optimized Root Cause Analysis
ORCA is an agent-orchestrated interactive copilot that automates and guides end-to-end causal analysis from workflow selection to report generation across real-world use cases.
-
CausalGuard: Conformal Inference under Graph Uncertainty
CausalGuard aggregates LLM-proposed and data-pruned DAGs to weight doubly robust pseudo-outcomes and applies conformal calibration to deliver finite-sample marginal coverage for conditional average treatment effects under graph uncertainty.
-
Enhancing Regime Shift Detection Using Unstructured Data: A Study on the Treasury Market
Using FOMC minutes to propose regime-shift candidates and a lenient text check to ratify data-detected candidates, the pipeline reaches F1=0.82 on 26 monetary-policy anchors, beating every data-only baseline.
-
CIVeX: Causal Intervention Verification for Language Agents
CIVeX maps agent tool calls to structural causal queries, checks identifiability, and issues auditable verdicts to prevent false executions while preserving utility on confounded benchmarks.
-
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
Four axioms (Causality, Minimality, Separability, Stability) are formalized for latent thought representations; audits of open LLMs on 23 tasks show none satisfy all four and representations add little beyond input embeddings.
-
Thinking Fast, Thinking Wrong: Intuitiveness Modulates LLM Counterfactual Reasoning in Policy Evaluation
Chain-of-thought helps LLMs on obvious policy findings but its benefit collapses on counter-intuitive ones, with case intuitiveness dominating model and prompt choice.
-
Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment
Introduces the CAUSALT3 benchmark for causal reasoning across Pearl's ladder and Regulated Causal Anchoring (RCA) to reduce sycophancy and skepticism in LLMs via inference-time verification.
-
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
Introduces CounterBench benchmark and CoIn iterative reasoning method showing LLMs perform near random on formal counterfactual tasks but improve substantially with guided backtracking.
-
CasualSynth: Generating Structurally Sound Synthetic Data
CausalSynth combines structural causal models with LLMs and iterative verification to produce synthetic data that respects given causal structures while remaining linguistically natural.
-
Hume's Representational Conditions for Causal Judgment: What Bayesian Formalization Abstracted Away
Hume's causal judgment requires experiential grounding, structured retrieval, and vivacity transfer, conditions that Bayesian formalizations abstract away while LLMs retain only statistical updating.
-
The New Associationism: Lessons from Deep Learning
Supervised learning across AI systems vindicates a uniform error-driven associationism for cognition, though operating inside advanced computational structures beyond classical associationist models.
-
Large Language Models for Causal Relations Extraction in Social Media: A Validation Framework for Disaster Intelligence
The authors introduce a validation framework showing LLMs can pull causal links from disaster social media but require checks against post-event evidence to avoid relying on model priors.
-
Gemma 3 Technical Report
Gemma 3 introduces multimodal open models with architectural changes for efficient long context, trained via distillation and a new post-training recipe that makes the 4B version competitive with prior 27B models and the 27B version comparable to Gemini-1.5-Pro.
-
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
ODYSSEY is a sheaf-theoretic framework for building verifiable foundation models as compositions of foundries via left and right Kan extensions.
- DeepImagine: Clinical Trial Outcome Prediction via Stepwise Local Counterfactual Imaginations