Pith. sign in

REVIEW 1 cited by

JSONoid: Monoid-based Enrichment for Configurable and Scalable Data-Driven Schema Discovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.03113 v1 pith:PJML2O44 submitted 2023-07-06 cs.DB

classification cs.DB
keywords datadiscoverydistributedschemajsonoidadditionalapproachesexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Schema discovery is an important aspect to working with data in formats such as JSON. Unlike relational databases, JSON data sets often do not have associated structural information. Consumers of such datasets are often left to browse through data in an attempt to observe commonalities in structure across documents to construct suitable code for data processing. However, this process is time-consuming and error-prone. Existing distributed approaches to mining schemas present a significant usability advantage as they provide useful metadata for large data sources. However, depending on the data source, ad hoc queries for estimating other properties to help with crafting an efficient data pipeline can be expensive. We propose JSONoid, a distributed schema discovery process augmented with additional metadata in the form of monoid data structures that are easily maintainable in a distributed setting. JSONoid subsumes several existing approaches to distributed schema discovery with similar performance. Our approach also adds significant useful additional information about data values to discovered schemas with linear scalability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Introducing Schema Inference as a Scalable SQL Function [Extended Version]

    cs.DB 2024-11 conditional novelty 6.0 of 10

    Schema inference is now a native SQL aggregate in Apache AsterixDB, using local schema trees and global merging, and benchmarks show speedups of up to two orders of magnitude over Spark-based tools.

Pith tools