{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:MUDJY7K5QSTNU3T2D4YH57ZTYM","short_pith_number":"pith:MUDJY7K5","schema_version":"1.0","canonical_sha256":"65069c7d5d84a6da6e7a1f307eff33c3391e702504a8a1c083bf936a480f9f31","source":{"kind":"arxiv","id":"2511.01680","version":4},"attestation_state":"computed","paper":{"title":"Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach","license":"http://creativecommons.org/licenses/by/4.0/","headline":"A framework maps unstructured data to concept embeddings and uses selective inference to produce statistically valid interpretable discoveries.","cross_cats":["cs.LG"],"primary_cat":"econ.EM","authors_text":"Jacob Carlson","submitted_at":"2025-11-03T15:42:32Z","abstract_excerpt":"Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) manually pre-specify all important aspects of the unstructured data to measure; they are interested in \"discovery.\" This paper proposes a general and flexible framework for pursuing such discovery from unstructured data in a statistically principled way. "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":true},"canonical_record":{"source":{"id":"2511.01680","kind":"arxiv","version":4},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"econ.EM","submitted_at":"2025-11-03T15:42:32Z","cross_cats_sorted":["cs.LG"],"title_canon_sha256":"8c22c963fc5fcc58abde3ed5d6595a18bcec7f82c8c499b9540f710fb6820005","abstract_canon_sha256":"16b33d4519d38679a1f8c7cda4ecb4e47be1d7b46f0752a6916395bd3f30e32e"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-16T01:22:29.352445Z","signature_b64":"YkBQZtgQ5KHTUlWZYoXa7z9wpgcKYS8oWd/i0sEKrD71gPmUnwMB43mxvi6gimfFTiA9LDhH1JytpkS+zuNcAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"65069c7d5d84a6da6e7a1f307eff33c3391e702504a8a1c083bf936a480f9f31","last_reissued_at":"2026-07-16T01:22:29.351556Z","signature_status":"signed_v1","first_computed_at":"2026-07-16T01:22:29.351556Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach","license":"http://creativecommons.org/licenses/by/4.0/","headline":"A framework maps unstructured data to concept embeddings and uses selective inference to produce statistically valid interpretable discoveries.","cross_cats":["cs.LG"],"primary_cat":"econ.EM","authors_text":"Jacob Carlson","submitted_at":"2025-11-03T15:42:32Z","abstract_excerpt":"Social scientists are increasingly turning to unstructured datasets to unlock new empirical insights, e.g., estimating descriptive statistics of or causal effects on quantitative measures derived from text, audio, or video data. In many settings, unsupervised analysis is of primary interest, in that the researcher does not want to (or cannot) manually pre-specify all important aspects of the unstructured data to measure; they are interested in \"discovery.\" This paper proposes a general and flexible framework for pursuing such discovery from unstructured data in a statistically principled way. "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"The framework leverages recent methods from the literature on AI interpretability to map unstructured data points to high-dimensional, sparse, and interpretable 'concept embeddings'; computes statistics from these concept embeddings for testing interpretable, concept-by-concept hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ('discoveries'); and both generates and evaluates human-interpretable natural language descriptions of these discoveries.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The selective inference procedures remain valid when applied to statistics derived from AI-generated concept embeddings rather than from pre-specified variables; this relies on the new high-dimensional central limit theory results holding for the particular dependence structure induced by the embedding step (abstract, paragraph describing the framework).","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"A framework maps unstructured data to concept embeddings and uses selective inference to produce statistically valid interpretable discoveries.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"bbe161e036cb6234d0c365bdde9da231eabb8ff4b9f4a4e629a374a1cc9b1827"},"source":{"id":"2511.01680","kind":"arxiv","version":4},"verdict":{"id":"cf074293-33d9-458f-b520-f25c0523b4a5","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-18T01:54:06.637306Z","strongest_claim":"The framework leverages recent methods from the literature on AI interpretability to map unstructured data points to high-dimensional, sparse, and interpretable 'concept embeddings'; computes statistics from these concept embeddings for testing interpretable, concept-by-concept hypotheses; performs selective inference on these hypotheses using algorithms validated by new results in high-dimensional central limit theory, producing a selected set ('discoveries'); and both generates and evaluates human-interpretable natural language descriptions of these discoveries.","one_line_summary":"A new framework combines AI-derived concept embeddings with high-dimensional selective inference to enable statistically principled, interpretable discovery from unstructured data in empirical economics.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The selective inference procedures remain valid when applied to statistics derived from AI-generated concept embeddings rather than from pre-specified variables; this relies on the new high-dimensional central limit theory results holding for the particular dependence structure induced by the embedding step (abstract, paragraph describing the framework).","pith_extraction_headline":"A framework maps unstructured data to concept embeddings and uses selective inference to produce statistically valid interpretable discoveries."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2511.01680/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":1,"snapshot_sha256":"bc9b423fb20fe991c6ce644da9b7b6920c686ec1531417885a57b9cba536f891"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2511.01680","created_at":"2026-07-16T01:22:29.351967+00:00"},{"alias_kind":"arxiv_version","alias_value":"2511.01680v4","created_at":"2026-07-16T01:22:29.351967+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2511.01680","created_at":"2026-07-16T01:22:29.351967+00:00"},{"alias_kind":"pith_short_12","alias_value":"MUDJY7K5QSTN","created_at":"2026-07-16T01:22:29.351967+00:00"},{"alias_kind":"pith_short_16","alias_value":"MUDJY7K5QSTNU3T2","created_at":"2026-07-16T01:22:29.351967+00:00"},{"alias_kind":"pith_short_8","alias_value":"MUDJY7K5","created_at":"2026-07-16T01:22:29.351967+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":1,"internal_anchor_count":1,"sample":[{"citing_arxiv_id":"2603.26930","citing_title":"In your own words: computationally identifying interpretable themes in free-text survey data","ref_index":113,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":1,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM","json":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM.json","graph_json":"https://pith.science/api/pith-number/MUDJY7K5QSTNU3T2D4YH57ZTYM/graph.json","events_json":"https://pith.science/api/pith-number/MUDJY7K5QSTNU3T2D4YH57ZTYM/events.json","paper":"https://pith.science/paper/MUDJY7K5"},"agent_actions":{"view_html":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM","download_json":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM.json","view_paper":"https://pith.science/paper/MUDJY7K5","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2511.01680&json=true","fetch_graph":"https://pith.science/api/pith-number/MUDJY7K5QSTNU3T2D4YH57ZTYM/graph.json","fetch_events":"https://pith.science/api/pith-number/MUDJY7K5QSTNU3T2D4YH57ZTYM/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM/action/timestamp_anchor","attest_storage":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM/action/storage_attestation","attest_author":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM/action/author_attestation","sign_citation":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM/action/citation_signature","submit_replication":"https://pith.science/pith/MUDJY7K5QSTNU3T2D4YH57ZTYM/action/replication_record"}},"created_at":"2026-07-16T01:22:29.351967+00:00","updated_at":"2026-07-16T01:22:29.351967+00:00"}