The foundational anatomy of machine learning datasets: decomposing tabular data into input feature matrices (X) and target vectors (y).
# Extracting Features (X) and Labels (y) in Python using Pandas# ───────────────────────────────────────────────────────────────import pandas as pdfrom sklearn.linear_model import LinearRegression # 1. Create a sample housing dataset DataFrameraw_data = { "SquareFeet": [850, 1200, 1500, 2100, 2800], "Bedrooms": [1, 2, 3, 3, 4], "Bathrooms": [1.0, 1.5, 2.0, 2.5, 3.0], "HasGarage": [0, 1, 1, 1, 1], "Price": [195000, 260000, 315000, 425000, 580000] # TARGET LABEL}df = pd.DataFrame(raw_data)print("Original Dataset:\n", df) # 2. Separate into Feature Matrix (X) and Target Vector (y)# Feature Matrix X: Drop the target label columnX = df.drop(columns=["Price"]) # Target Vector y: Extract the target label seriesy = df["Price"] print(f"\nFeature Matrix X Shape: {X.shape} (5 samples, 4 features)")print(f"Target Vector y Shape: {y.shape} (5 labels)") # 3. Train Model on (X, y)model = LinearRegression()model.fit(X, y) # 4. Inference on a new house (X_new has features, but NO label!)X_new = pd.DataFrame({ "SquareFeet": [1800], "Bedrooms": [3], "Bathrooms": [2.0], "HasGarage": [1]})predicted_price = model.predict(X_new)[0]print(f"\nPredicted Price for New House: ${predicted_price:,.2f}")Inspect real-world tabular datasets, examine sample vectors, and simulate production inference.
| Sample | FEATURE MATRIX (X) — Input Variables | TARGET LABEL (y) — Ground Truth | ||||
|---|---|---|---|---|---|---|
| # (m) | X₁SquareFeet | X₂Bedrooms | X₃Bathrooms | X₄ZipCodeScore | X₅HasGarage | yPrice ($) |
| Row 1 | 850 | 1 | 1 | 6.2 | 0 | $195,000 |
| Row 2 | 1200 | 2 | 1.5 | 7.5 | 1 | $260,000 |
| Row 3 | 1550 | 3 | 2 | 8.1 | 1 | $335,000 |
| Row 4 | 2100 | 3 | 2.5 | 8.8 | 1 | $445,000 |
| Row 5 | 2800 | 4 | 3 | 9.4 | 1 | $595,000 |
Mathematical representation of observation row #1:
New unseen instance has features, but NO label (y is unknown):
In a machine learning project to predict employee salary, a dataset has columns: [Age, YearsOfExperience, EducationLevel, Department, AnnualSalary]. Which of the following correctly describes the Features (X) and Label (y)?