Pith. sign in
Pith Number

pith:MUDJY7K5

pith:2025:MUDJY7K5QSTNU3T2D4YH57ZTYM
not attested not anchored not stored refs pending

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach

Jacob Carlson

A framework maps unstructured data to concept embeddings and uses selective inference to produce statistically valid interpretable discoveries.

arxiv:2511.01680 v4 · 2025-11-03 · econ.EM · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{MUDJY7K5QSTNU3T2D4YH57ZTYM}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

The framework leverages recent methods from the literature on AI interpretability to map unstructured data points to high-dimensional, sparse, and interpretable 'concept embeddings'; computes statistics from these concept embeddings for testing interpretable, concept-by-concept hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ('discoveries'); and both generates and evaluates human-interpretable natural language descriptions of these discoveries.

C2weakest assumption

The selective inference procedures remain valid when applied to statistics derived from AI-generated concept embeddings rather than from pre-specified variables; this relies on the new high-dimensional central limit theory results holding for the particular dependence structure induced by the embedding step (abstract, paragraph describing the framework).

C3one line summary

A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.

Formal links

1 machine-checked theorem link

Cited by

1 paper in Pith

Receipt and verification
First computed 2026-07-16T01:22:29.351556Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

65069c7d5d84a6da6e7a1f307eff33c3391e702504a8a1c083bf936a480f9f31

Aliases

arxiv: 2511.01680 · arxiv_version: 2511.01680v4 · doi: 10.48550/arxiv.2511.01680 · pith_short_12: MUDJY7K5QSTN · pith_short_16: MUDJY7K5QSTNU3T2 · pith_short_8: MUDJY7K5
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 65069c7d5d84a6da6e7a1f307eff33c3391e702504a8a1c083bf936a480f9f31
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "16b33d4519d38679a1f8c7cda4ecb4e47be1d7b46f0752a6916395bd3f30e32e",
    "cross_cats_sorted": [
      "cs.LG"
    ],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "econ.EM",
    "submitted_at": "2025-11-03T15:42:32Z",
    "title_canon_sha256": "8c22c963fc5fcc58abde3ed5d6595a18bcec7f82c8c499b9540f710fb6820005"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2511.01680",
    "kind": "arxiv",
    "version": 4
  }
}