Understanding the foundational split: learning with an explicit ground-truth supervisor vs discovering hidden patterns autonomously.
# Comparing Supervised Learning vs Unsupervised Learning in Python# ───────────────────────────────────────────────────────────────────import numpy as npfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.cluster import KMeans # 1. SUPERVISED LEARNING: Training with Input Features (X) + Target Labels (y)# Data: [Weight in grams, Smoothness score 1-10]X_supervised = np.array([ [150, 8], # Apple [170, 7], # Apple [130, 9], # Apple [200, 3], # Pear [220, 2], # Pear [210, 4] # Pear])# Supervisor provides explicit ground-truth class labels: 0 = Apple, 1 = Peary_supervised = np.array([0, 0, 0, 1, 1, 1]) clf = RandomForestClassifier(n_estimators=10, random_state=42)clf.fit(X_supervised, y_supervised) # Learns mapping X -> y new_fruit = np.array([[160, 8.5]])pred_class = clf.predict(new_fruit)[0]print(f"Supervised Prediction: {'Apple' if pred_class == 0 else 'Pear'}") # 2. UNSUPERVISED LEARNING: Training with ONLY Features (X) — NO Labels!# The algorithm receives identical measurements without any names or labelsX_unsupervised = np.array([ [150, 8], [170, 7], [130, 9], [200, 3], [220, 2], [210, 4]]) # K-Means autonomously groups data points into k=2 natural geometric clusterskmeans = KMeans(n_clusters=2, random_state=42)cluster_labels = kmeans.fit_predict(X_unsupervised) print(f"Unsupervised Discovered Clusters: {cluster_labels}")# Result: [0, 0, 0, 1, 1, 1] — Discovered the exact two fruit groups without labels!Interactive comparison of learning with ground-truth supervisor labels vs autonomous unsupervised pattern discovery.
A retail eCommerce company has transaction records for 2,000,000 customers with purchase histories, but NO predetermined customer segment categories. They want to group customers into 5 spending personas for targeted marketing. Which approach should they use?