Clustering Performance on Evolving Data Streams: Assessing Algorithms and Evaluation Measures within MOA

In today's applications, evolving data streams are ubiquitous. Stream clustering algorithms were introduced to gain useful knowledge from these streams in real-time. The quality of the obtained clusterings, i.e. how good they reflect the data, can be assessed by evaluation measures. A multitude of stream clustering algorithms and evaluation measures for clusterings were introduced in the literature; however, until now there is no general tool for a direct comparison of the different algorithms or the evaluation measures. In our demo, we present a novel experimental framework for both tasks.  It offers the means for extensive evaluation and visualization and is an extension of the Massive Online Analysis (MOA) software environment released under the GNU GPL License.

Authors: Kranen P., Kremer H., Jansen T., Seidl T., Bifet A., Holmes G., Pfahringer B.
Published in: Proc. IEEE International Conference on Data Mining (ICDM 2010), Sydney, Australia
Publisher: IEEE Computer Society - Washington,USA
Sprache: EN
Jahr: 2010


Seiten: 1400-1403
ISBN: 978-1-4244-9244-2
Konferenz: ICDM
Typ: Tagungsbeiträge
Forschungsgebiet: Data Analysis and Knowledge Extraction