Pith. sign in

REVIEW 1 cited by

Uncertain Centroid based Partitional Clustering of Uncertain Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1203.6401 v1 pith:ZR5JCXPK submitted 2012-03-29 cs.DB

classification cs.DB
keywords clusteringuncertaincentroidclusterdatapartitionalbeenexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Clustering uncertain data has emerged as a challenging task in uncertain data management and mining. Thanks to a computational complexity advantage over other clustering paradigms, partitional clustering has been particularly studied and a number of algorithms have been developed. While existing proposals differ mainly in the notions of cluster centroid and clustering objective function, little attention has been given to an analysis of their characteristics and limits. In this work, we theoretically investigate major existing methods of partitional clustering, and alternatively propose a well-founded approach to clustering uncertain data based on a novel notion of cluster centroid. A cluster centroid is seen as an uncertain object defined in terms of a random variable whose realizations are derived based on all deterministic representations of the objects to be clustered. As demonstrated theoretically and experimentally, this allows for better representing a cluster of uncertain objects, thus supporting a consistently improved clustering performance while maintaining comparable efficiency with existing partitional clustering algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Stochastic SketchRefine: Scaling In-Database Decision-Making under Uncertainty to Millions of Tuples

    cs.DB 2024-11 conditional novelty 7.0 of 10

    A new linearization and partitioning framework lets stochastic package queries with value-at-risk or conditional-value-at-risk constraints run on millions of tuples in minutes.

Pith tools