Skip to content

Latest commit

ย 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

Iris Dataset Clustering using K-Means

Project Overview

This project explores unsupervised learning using the K-Means clustering algorithm on the Iris dataset.

The primary goal of this project was to understand the complete workflow of clustering, including:

  • How K-Means forms clusters
  • The role of centroid initialization
  • Understanding inertia
  • Evaluating clustering quality using Silhouette Score
  • Analyzing the effect of PCA on clustering performance

The implementation was intentionally done step-by-step to build strong conceptual understanding of K-Means rather than using high-level abstractions.


๐Ÿ“Š Dataset

This project uses the built-in Iris dataset provided by Scikit-learn.

Dataset Characteristics

  • 150 samples
  • 4 numerical features:
    • Sepal Length
    • Sepal Width
    • Petal Length
    • Petal Width
  • 3 species (used only for reference; clustering itself is unsupervised)

Dataset Loading

from sklearn.datasets import load_iris

๐Ÿ“Š Visual Results

๐Ÿ”น K-Means Cluster Visualization

KMeans Cluster


๐Ÿ”น PCA Visualization

PCA Visualization


๐Ÿ”น Elbow Method

Elbow Plot

About

Experimental analysis of K-Means clustering on the Iris dataset, including Elbow Method, Silhouette Score evaluation, and comparison of PCA-reduced vs full-dimensional feature space.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages