REVIEW 3 cited by
Using GANs for Sharing Networked Time Series Data: Challenges, Initial Promise, and Open Questions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Limited data access is a longstanding barrier to data-driven research and development in the networked systems community. In this work, we explore if and how generative adversarial networks (GANs) can be used to incentivize data sharing by enabling a generic framework for sharing synthetic datasets with minimal expert knowledge. As a specific target, our focus in this paper is on time series datasets with metadata (e.g., packet loss rate measurements with corresponding ISPs). We identify key challenges of existing GAN approaches for such workloads with respect to fidelity (e.g., long-term dependencies, complex multidimensional relationships, mode collapse) and privacy (i.e., existing guarantees are poorly understood and can sacrifice fidelity). To improve fidelity, we design a custom workflow called DoppelGANger (DG) and demonstrate that across diverse real-world datasets (e.g., bandwidth measurements, cluster requests, web sessions) and use cases (e.g., structural characterization, predictive modeling, algorithm comparison), DG achieves up to 43% better fidelity than baseline models. Although we do not resolve the privacy problem in this work, we identify fundamental challenges with both classical notions of privacy and recent advances to improve the privacy properties of GANs, and suggest a potential roadmap for addressing these challenges. By shedding light on the promise and challenges, we hope our work can rekindle the conversation on workflows for data sharing.
Forward citations
Cited by 3 Pith papers
-
Synthetic Time Series Generation via Complex Networks
A first-order Markov chain on empirical quantile bins — built from the original series and sampled to produce new series — reproduces marginal distributions and lag-1 correlations but cannot capture long-range or high...
-
Federated Diffusion Modeling with Differential Privacy for Tabular Data Synthesis
DP-FedTabDiff wraps an existing federated tabular diffusion model with per-client DP-SGD and reports how the privacy budget, number of clients, and local update count affect synthetic data quality and empirical privacy risk.
-
Synthetic Time Series Data Generation for Healthcare Applications: A PCG Case Study
A comparison of three generative models for synthetic PCG signals reports low distribution-distance scores, but the evaluation lacks held-out splits, baselines, and error bars.
Discussion (0). Continue with ORCID to comment.