Pith. sign in

REVIEW 3 cited by

Large Language Models for JSON Schema Discovery

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.03286 v1 pith:3PBDHBAM submitted 2024-07-03 cs.DB

classification cs.DB
keywords datajsonschemainformationlanguagemodelsschemasuseful
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Semi-structured data formats such as JSON have proved to be useful data models for applications that require flexibility in the format of data stored. However, JSON data often come without the schemas that are typically available with relational data. This has resulted in a number of tools for discovering schemas from a collection of data. Although such tools can be useful, existing approaches focus on the syntax of documents and ignore semantic information. In this work, we explore the automatic addition of meaningful semantic information to discovered schemas similar to information that is added by human schema authors. We leverage large language models and a corpus of manually authored JSON Schema documents to generate natural language descriptions of schema elements, meaningful names for reusable definitions, and identify which discovered properties are most useful and which can be considered "noise". Our approach performs well on existing metrics for text generation that have been previously shown to correlate well with human judgement.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

    cs.DL 2026-07 accept novelty 5.5 of 10

    Sixteen expert-annotated scientific-process schemas, built via LLM-assisted human-in-the-loop mining and released in JSON Schema and SHACL, form the first SciSchema.org collection.

  2. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  3. Towards Next Generation Data Engineering Pipelines

    cs.DB 2025-07 unverdicted novelty 5.0 of 10

    A vision paper defines three levels of next generation data engineering pipelines (optimized, self-aware, self-adapting) and proposes an architecture to realize them.

Pith tools