Pith. sign in

REVIEW 2 cited by

Sequential Models in the Synthetic Data Vault

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.14406 v1 pith:CT52MQMC submitted 2022-07-28 cs.LG cs.MS

classification cs.LGcs.MS
keywords datasyntheticmodelsequentialmodelsvaultcalledgenerative
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The goal of this paper is to describe a system for generating synthetic sequential data within the Synthetic data vault. To achieve this, we present the Sequential model currently in SDV, an end-to-end framework that builds a generative model for multi-sequence, real-world data. This includes a novel neural network-based machine learning model, conditional probabilistic auto-regressive (CPAR) model. The overall system and the model is available in the open source Synthetic Data Vault (SDV) library {https://github.com/sdv-dev/SDV}, along with a variety of other models for different synthetic data needs. After building the Sequential SDV, we used it to generate synthetic data and compared its quality against an existing, non-sequential generative adversarial network based model called CTGAN. To compare the sequential synthetic data against its real counterpart, we invented a new metric called Multi-Sequence Aggregate Similarity (MSAS). We used it to conclude that our Sequential SDV model learns higher level patterns than non-sequential models without any trade-offs in synthetic data quality.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Generative Models Keep Time? A Time-Aware Evaluation of Synthetic Sequential Tabular Data

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Static-distribution fidelity is a poor proxy for temporal fidelity in synthetic sequential tabular data; measuring timestamp, trajectory, cross-sectional, and relational structure over time changes model rankings.

  2. Resampling Methods that Generate Time Series Data to Enable Sensitivity and Model Analysis in Energy Modeling

    stat.CO 2025-02 conditional novelty 4.0 of 10

    Two non-parametric bootstrap schemes plus two displacement methods can produce synthetic energy time series that resemble the original series statistically, but the validity claim rests on descriptive statistics and t...

Pith tools