Alerte : Maintenance en cours. Certains ouvrages sont temporairement indisponibles et reviendront bientôt.

Clustering for Data Mining: A Data Recovery Approach

Book Details
Title Clustering for Data Mining: A Data Recovery Approach
Author(s) Boris Mirkin
Publisher Chapman & Hall
Year 2005
Edition 1st Edition
Language English
Pages 260 pages
ISBN 9781584885344
Genre / Domain Computer Science, Data Science, Machine Learning
Series Unknown
Size 3.88 MB
Extension PDF

Summary

Clustering for Data Mining: A Data Recovery Approach, authored by Boris Mirkin and published by Chapman & Hall in 2005, presents a groundbreaking shift in the way clustering is understood and applied. Traditionally viewed as more of an art than a science, clustering has often relied on ad hoc techniques and trial-and-error. This book challenges that paradigm by introducing a rigorous theoretical framework—the data recovery approach—which transforms clustering into a well-founded methodology grounded in statistical and mathematical principles .

The book systematically develops a theory that bridges the gaps between the popular K-Means partitioning method and Ward's hierarchical clustering method, demonstrating their underlying connections and unifying them within a single conceptual framework . This approach is extended to cover contemporary challenges in data mining, including the clustering of mixed-scale data (combining continuous and categorical variables) and handling incomplete clustering scenarios . The author also discusses related areas such as principal component analysis, contingency measures, and data visualization, providing a holistic view of exploratory data analysis .

What makes this book particularly valuable is its practical, hands-on focus. It includes nearly 60 computational examples that walk readers through all stages of the clustering process—from data preprocessing and validation to the interpretation of results . The examples, which are based on both real-world and synthetic datasets, are designed to help readers understand how to apply the theory in practice. This makes the book an excellent resource for both learning and professional application, as it provides concrete, reproducible guidance for conducting clustering analyses .

This book is well-suited for a wide range of audiences, including educators, researchers, and professionals in data mining, machine learning, statistics, and related fields. It can be used as a textbook for advanced undergraduate or graduate courses on cluster analysis and data mining, as well as a reference for practitioners seeking a deeper understanding of clustering methods . The theoretical rigor combined with practical examples ensures that readers gain both the conceptual foundation and the practical skills needed to apply clustering techniques effectively .

As a significant contribution to the field, Clustering for Data Mining has been recognized for its innovative approach and clarity. The author, a distinguished expert in data analysis and clustering, provides a unified perspective that has influenced both academic research and practical applications. The book's focus on data recovery not only improves the interpretability of clustering results but also opens new avenues for developing advanced clustering algorithms .

Key Features

  • Presents a novel data recovery theory for clustering that bridges K-Means and Ward's hierarchical methods.
  • Extends clustering techniques to handle mixed-scale data and incomplete clustering scenarios.
  • Includes nearly 60 computational examples covering all stages of clustering from preprocessing to interpretation.
  • Provides a unified framework connecting clustering with principal component analysis and data visualization.
  • Offers a rigorous theoretical foundation that transforms clustering from an art to a science.
  • Features real-world and synthetic datasets to illustrate practical applications.
  • Suitable for teaching, self-study, and professional reference in data mining and machine learning.
  • Written by a leading expert in data analysis with extensive research experience.
  • Addresses contemporary challenges such as mixed-scale data and incomplete clustering.
  • Includes detailed explanations of algorithms and their theoretical underpinnings.
  • Published by Chapman & Hall, a respected publisher in scientific and technical literature.
  • Provides a comprehensive bibliography for further reading and research.

About the Author

Boris Mirkin is a Professor in the Department of Computer Science and Information Systems at Birkbeck, University of London. He is a leading expert in the fields of data analysis, clustering, and data mining, with a particular focus on developing theoretical foundations for practical data analysis methods. He has held academic positions at universities in Russia, the UK, and the USA, and has published numerous research papers and books on data analysis, clustering, and related topics.

His research interests include intelligent data analysis, multivariate statistics, and the application of mathematical models to real-world problems. He is known for his work on the data recovery approach to clustering, which has provided a unified framework for understanding and applying clustering methods.

Related Books

  • Cluster Analysis — Brian S. Everitt, Sabine Landau, Morven Leese
  • Introduction to Data Mining — Pang-Ning Tan, Michael Steinbach, Vipin Kumar
  • Data Clustering: Algorithms and Applications — Charu C. Aggarwal, Chandan K. Reddy
  • Pattern Recognition and Machine Learning — Christopher M. Bishop
  • The Elements of Statistical Learning — Trevor Hastie, Robert Tibshirani, Jerome Friedman
  • Mining of Massive Datasets — Jure Leskovec, Anand Rajaraman, Jeffrey D. Ullman

Ads

Frequently Asked Questions

Q : What is the main contribution of this book to the field of clustering?

R : This book introduces a novel data recovery theory for clustering that provides a unified framework for understanding and applying clustering methods. It bridges the gaps between K-Means and Ward's hierarchical methods and extends them to handle contemporary challenges such as mixed-scale data and incomplete clustering .

Q : Who is the author and what is his background?

R : The author is Boris Mirkin, a Professor at Birkbeck, University of London. He is a leading expert in data analysis, clustering, and data mining, with extensive research experience and publications in the field .

Q : What topics are covered in this book?

R : The book covers the theory of data recovery clustering, K-Means and Ward's methods, clustering of mixed-scale data, incomplete clustering, principal component analysis, contingency measures, data visualization, and includes nearly 60 computational examples covering all stages of clustering .

Q : Is this book suitable for practitioners as well as researchers?

R : Yes, the book is designed to be useful for teaching, self-study, and professional reference. Its focus on data recovery methods, theory-based guidance, and practical instructions makes it valuable for both practitioners and researchers in data mining and machine learning .

Q : What is the edition and publication year of this book?

R : This is the first edition, published in 2005 by Chapman & Hall .

Q : How does this book differ from other clustering textbooks?

R : Unlike many clustering textbooks that focus on algorithmic techniques, this book provides a rigorous theoretical foundation based on data recovery theory. It offers a unified perspective that connects different clustering methods and addresses modern data mining challenges .

Q : Are there practical examples included in the book?

R : Yes, the book includes nearly 60 computational examples that cover all stages of clustering, from data preprocessing and validation to cluster interpretation, using both real-world and synthetic datasets .

Enregistrer un commentaire

Thanks for comment

Page précédente Accueil Page suivante

Post Share Buttons

Les plus populaires Voir la suite

Biblio-Sciences