Pith. sign in

REVIEW 1 cited by

Lightweight Knowledge Representations for Automating Data Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12848 v1 pith:UFGQ7N3W submitted 2023-10-15 cs.DB cs.AI

classification cs.DBcs.AI
keywords dataanalysisanalyticinformationknowledgetaxonomyanalyticsautomating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The principal goal of data science is to derive meaningful information from data. To do this, data scientists develop a space of analytic possibilities and from it reach their information goals by using their knowledge of the domain, the available data, the operations that can be performed on those data, the algorithms/models that are fed the data, and how all of these facets interweave. In this work, we take the first steps towards automating a key aspect of the data science pipeline: data analysis. We present an extensible taxonomy of data analytic operations that scopes across domains and data, as well as a method for codifying domain-specific knowledge that links this analytics taxonomy to actual data. We validate the functionality of our analytics taxonomy by implementing a system that leverages it, alongside domain labelings for 8 distinct domains, to automatically generate a space of answerable questions and associated analytic plans. In this way, we produce information spaces over data that enable complex analyses and search over this data and pave the way for fully automated data analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

    cs.DB 2026-08 conditional novelty 6.0 of 10

    A neurosymbolic pipeline automatically constructs executable analytic semantic schemas ("rings") from relational databases, claiming 100% coverage and retrieval pass rates across eight domains.

Pith tools