Pith. sign in

REVIEW 4 cited by

Causal Inference on Outcomes Learned from Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.00725 v1 pith:JXEUCXGS submitted 2025-03-02 econ.EM cs.CLcs.LGstat.ME

classification econ.EMcs.CLcs.LGstat.ME
keywords textcausalinferenceoutcomesacrossdocumentsgroupsneed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a machine-learning tool that yields causal inference on text in randomized trials. Based on a simple econometric framework in which text may capture outcomes of interest, our procedure addresses three questions: First, is the text affected by the treatment? Second, which outcomes is the effect on? And third, how complete is our description of causal effects? To answer all three questions, our approach uses large language models (LLMs) that suggest systematic differences across two groups of text documents and then provides valid inference based on costly validation. Specifically, we highlight the need for sample splitting to allow for statistical validation of LLM outputs, as well as the need for human labeling to validate substantive claims about how documents differ across groups. We illustrate the tool in a proof-of-concept application using abstracts of academic manuscripts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Econometrics with Pre-Trained Embeddings for Unstructured Data

    econ.EM 2026-07 accept novelty 7.0 of 10

    Pre-trained embeddings are valid in double machine learning when the target nuisance function lies in the span of the source-task representation; under that condition the downstream estimator can converge faster than ...

  2. Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

    econ.EM 2025-11 unverdicted novelty 7.0 of 10

    A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.

  3. E-LDA: Toward Interpretable LDA Topic Models with Strong Guarantees in Logarithmic Parallel Time

    cs.LG 2025-06 conditional novelty 6.0 of 10

    E-LDA recasts document-topic assignment as submodular maximization and solves it near-optimally in logarithmic parallel time.

  4. Advertising in AI systems: Society must be vigilant

    cs.AI 2025-05 conditional novelty 4.0 of 10

    Generative AI outputs will likely carry embedded commercial content, and the paper proposes design principles, provenance tracking, and two debiasing strategies to preserve transparency.

Pith tools