Pith. sign in

REVIEW 1 cited by

Generating Multidimensional Clusters With Support Lines

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.10327 v3 pith:L5VBO4PQ submitted 2023-01-24 cs.LG cs.CVcs.PL

classification cs.LGcs.CVcs.PL
keywords datasyntheticalgorithmsclugenclusteringclusterscreatinggeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthetic data is essential for assessing clustering techniques, complementing and extending real data, and allowing for more complete coverage of a given problem's space. In turn, synthetic data generators have the potential of creating vast amounts of data -- a crucial activity when real-world data is at premium -- while providing a well-understood generation procedure and an interpretable instrument for methodically investigating cluster analysis algorithms. Here, we present Clugen, a modular procedure for synthetic data generation, capable of creating multidimensional clusters supported by line segments using arbitrary distributions. Clugen is open source, comprehensively unit tested and documented, and is available for the Python, R, Julia, and MATLAB/Octave ecosystems. We demonstrate that our proposal can produce rich and varied results in various dimensions, is fit for use in the assessment of clustering algorithms, and has the potential to be a widely used framework in diverse clustering-related research tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSMentor: Enhancing Data Science Agents with Curriculum Learning and Online Knowledge Accumulation

    cs.AI 2025-05 conditional novelty 5.0 of 10

    Ordering data science problems easy-to-hard and accumulating their solutions in a memory buffer improves LLM agent pass rates on DSEval and QRData by up to 5.2%.

Pith tools