Lecture 1 / 30
Lecture 01 · Scikit-Learn Basics

Introduction to Scikit-Learn

Learn what Scikit-Learn is, why it is one of the most widely used machine learning libraries in Python, how it fits into the Python data science ecosystem, and how its simple API makes machine learning easier to build and understand.

What You Will Learn
  • What Scikit-Learn (sklearn) is
  • Why sklearn is widely used in machine learning
  • What problems sklearn solves
  • How to install sklearn
  • How to import sklearn modules
  • The sklearn API pattern: fit / predict / transform
  • The difference between features and labels
  • What estimators, transformers, and predictors are
  • How sklearn works with NumPy and Pandas
  • How to create your first machine learning model

01. What is Scikit-Learn?

Scikit-Learn (imported as sklearn) is a powerful open-source machine learning library for Python. It provides simple and efficient tools for data mining, data analysis, data preprocessing, machine learning, model evaluation, and building predictive systems.

Scikit-Learn is designed to make common machine learning tasks easier. Instead of implementing algorithms such as linear regression, logistic regression, decision trees, support vector machines, or k-nearest neighbors from scratch, developers can use tested implementations provided by the library.

It is built on top of important Python scientific computing libraries such as NumPy and SciPy, and it works extremely well with Pandas for handling datasets and Matplotlib for visualization.

Simple Definition

Scikit-Learn is a Python machine learning library that provides ready-to-use tools for preparing data, training models, making predictions, and evaluating machine learning systems.

02. Why Use Scikit-Learn?

Machine learning involves much more than simply choosing an algorithm. A complete machine learning workflow usually includes loading data, cleaning data, preprocessing features, splitting datasets, training models, evaluating results, tuning parameters, and making predictions.

Scikit-Learn provides tools for almost every traditional machine learning step in a consistent and easy-to-learn API.

Key Reasons
  • One consistent API for many machine learning algorithms
  • Comprehensive tools for classification, regression, clustering, and dimensionality reduction
  • Powerful preprocessing and feature transformation utilities
  • Model evaluation and cross-validation tools
  • Hyperparameter tuning capabilities
  • Free and open source
  • Excellent documentation and a large community
  • Works seamlessly with NumPy, Pandas, SciPy, and Matplotlib
  • Useful for both learning machine learning and developing real applications

03. What Can Scikit-Learn Do?

Scikit-Learn covers many of the most common tasks encountered in traditional machine learning. The library contains implementations of supervised and unsupervised learning algorithms as well as tools for preparing and evaluating data.

Classification

Predict which category something belongs to. Examples include spam detection, customer churn prediction, sentiment classification, and image classification.

Regression

Predict a continuous numerical value. Examples include house price prediction, sales prediction, temperature prediction, and demand forecasting.

Clustering

Group similar data points without predefined labels. Examples include customer segmentation, document grouping, and pattern discovery.

Preprocessing

Scale numerical features, encode categorical values, handle transformations, and prepare raw data for machine learning algorithms.

Model Evaluation

Measure model performance using accuracy, precision, recall, F1-score, mean squared error, cross-validation, confusion matrices, and other metrics.

Dimensionality Reduction

Reduce the number of input features while attempting to preserve useful information using techniques such as PCA.

04. Installing Scikit-Learn

Scikit-Learn can be installed using Python's package manager, pip. It is normally installed inside the Python environment where you plan to run your machine learning programs.

Terminal
pip install scikit-learn numpy pandas matplotlib

The package is installed using the name scikit-learn, but it is imported in Python using the name sklearn.

Important Difference

Installation name: scikit-learn
Python import name: sklearn

If you are working inside a virtual environment, activate that environment before installing the package. This helps keep project dependencies isolated from other Python projects.

05. Importing Scikit-Learn

You normally do not need to import the entire Scikit-Learn library manually. Instead, import the specific class or function required for your machine learning task.

python
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
import sklearn

print(sklearn.**version**)

This approach keeps your code organized and makes it clear which machine learning components your program is using.

Scikit-Learn is divided into different modules. For example, sklearn.linear_model contains linear models, sklearn.preprocessing contains preprocessing tools, and sklearn.model_selection contains tools for splitting data and performing model selection.

06. The sklearn API Pattern

One of the greatest strengths of Scikit-Learn is its consistent API. Many algorithms use similar methods even though they solve completely different machine learning problems.

The most important methods you will encounter are fit(), predict(), and transform().

sklearn_pattern.py
from sklearn.linear_model import LinearRegression

# Step 1: Create the model

model = LinearRegression()

# Step 2: Train the model on training data

model.fit(X_train, y_train)

# Step 3: Make predictions on new data

predictions = model.predict(X_test)

print(predictions)
The Three Core Methods
  • fit(X, y) — learns patterns or parameters from training data
  • predict(X) — uses a trained model to generate predictions
  • transform(X) — changes or prepares data using a transformer

The fit() method is particularly important because this is where the machine learning algorithm learns from the training data. After fitting, the model can usually be used with predict() to make predictions on previously unseen data.

07. Features and Labels

Machine learning models usually work with two important types of information: features and labels.

Features are the input variables used by a model. A label, also called a target, is the value the model is trying to predict.

features_labels.py
import numpy as np

# Features: house size in square feet

X = np.array([
[1000],
[1500],
[2000],
[2500]
])

# Labels: house prices

y = np.array([
150000,
220000,
300000,
360000
])

Here, X contains the input feature, which is the size of the house. The array y contains the target values, which are the corresponding house prices.

Remember
  • X usually represents input features.
  • y usually represents the target or label.
  • The model learns a relationship between X and y during training.
  • New values of X can then be supplied to generate predictions.

08. Estimators, Transformers, and Predictors

Scikit-Learn uses a few important concepts to organize its machine learning tools.

Important Concepts
  • Estimator — an sklearn object that learns parameters from data using fit().
  • Predictor — an estimator that can make predictions using predict().
  • Transformer — an object that modifies data using methods such as transform().
  • Model — commonly refers to a trained machine learning estimator.

For example, LinearRegression is an estimator and predictor because it can learn from data and then make numerical predictions. StandardScaler is a transformer because it learns scaling parameters and uses them to transform features.

09. A Complete First sklearn Example

Let's combine the basic concepts and build a small linear regression model. The example uses house size as the feature and house price as the target.

first_model.py
import numpy as np
from sklearn.linear_model import LinearRegression

# Training data

X = np.array([
[1000],
[1500],
[2000],
[2500]
])

y = np.array([
150000,
220000,
300000,
360000
])

# Create the model

model = LinearRegression()

# Train the model

model.fit(X, y)

# Predict the price of a 1800 square foot house

prediction = model.predict([[1800]])

print("Predicted price:", prediction[0])

This example demonstrates the fundamental sklearn workflow: prepare data, create an estimator, call fit(), and then call predict().

Learning Tip

Do not focus on memorizing every machine learning algorithm at the beginning. First understand the common sklearn workflow. Once you understand fit(), predict(), preprocessing, and evaluation, learning individual algorithms becomes much easier.

10. Scikit-Learn and NumPy

Scikit-Learn works closely with NumPy. Machine learning datasets are commonly represented using NumPy arrays or similar array-like structures.

numpy_sklearn.py
import numpy as np
from sklearn.linear_model import LinearRegression

X = np.array([
[1],
[2],
[3],
[4]
])

y = np.array([
2,
4,
6,
8
])

model = LinearRegression()
model.fit(X, y)

print(model.predict([[5]]))

NumPy provides the numerical array structure, while Scikit-Learn provides the machine learning algorithms and workflow.

11. Scikit-Learn and Pandas

Pandas is commonly used to load, clean, inspect, and prepare datasets before sending them to Scikit-Learn.

pandas_sklearn.py
import pandas as pd
from sklearn.linear_model import LinearRegression

data = pd.DataFrame({
"hours": [1, 2, 3, 4, 5],
"scores": [45, 50, 60, 70, 80]
})

X = data[["hours"]]
y = data["scores"]

model = LinearRegression()
model.fit(X, y)

print(model.predict([[6]]))

This is a common real-world workflow. Pandas manages the dataset, while Scikit-Learn handles the machine learning model.

Typical Data Science Workflow

Pandas → load and clean data → NumPy / Pandas → prepare features → Scikit-Learn → train and evaluate model → Matplotlib → visualize results.

12. Training and Testing Data

A machine learning model should not normally be evaluated only on the same data used to train it. Instead, a dataset is commonly divided into training and testing portions.

The training data is used to learn patterns. The testing data is kept separate and used to estimate how well the trained model performs on unseen examples.

train_test.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))
Important

Never evaluate a model only on the same examples it was trained on when you want to understand how it generalizes to unseen data. Separate evaluation data helps provide a more realistic measurement of model performance.

13. The Machine Learning Workflow

Scikit-Learn is most useful when it is understood as part of a complete machine learning workflow rather than simply a collection of algorithms.

workflow
Collect Data
     ↓
Explore Data
     ↓
Clean Data
     ↓
Prepare Features
     ↓
Split Training / Testing Data
     ↓
Preprocess Data
     ↓
Create Model
     ↓
Train with fit()
     ↓
Predict with predict()
     ↓
Evaluate Model
     ↓
Improve and Tune Model

Each stage is important. A powerful algorithm cannot compensate for badly prepared data, data leakage, incorrect evaluation, or poorly selected features.

14. Common Scikit-Learn Module Categories

Scikit-Learn contains many modules. Learning how these modules are organized makes it easier to find the right tool.

Useful Modules
  • sklearn.linear_model — linear and logistic regression models
  • sklearn.tree — decision tree algorithms
  • sklearn.ensemble — ensemble algorithms such as random forests
  • sklearn.neighbors — nearest-neighbor algorithms
  • sklearn.svm — support vector machine algorithms
  • sklearn.cluster — clustering algorithms
  • sklearn.preprocessing — scaling, encoding, and feature transformations
  • sklearn.model_selection — train/test splitting and model selection
  • sklearn.metrics — model evaluation metrics
  • sklearn.decomposition — dimensionality reduction techniques

15. Common Beginner Mistakes

Beginners often make mistakes when learning machine learning libraries. Understanding these mistakes early will make your sklearn programs more reliable.

Avoid These Mistakes
  • Forgetting to install Scikit-Learn in the correct Python environment
  • Using scikit-learn instead of sklearn in an import statement
  • Training and evaluating a model on exactly the same data
  • Forgetting to preprocess features when the algorithm requires it
  • Ignoring missing values or invalid data
  • Using inappropriate evaluation metrics
  • Assuming a more complicated model is automatically better
  • Ignoring data leakage during preprocessing and evaluation

Machine learning is not simply about obtaining a prediction. You must also understand how the data was prepared, how the model was trained, and how its performance was measured.

16. Key Points

Key Points
  • Scikit-Learn is a major Python library for traditional machine learning.
  • Install it using pip install scikit-learn.
  • The package is installed as scikit-learn but imported as sklearn.
  • Scikit-Learn provides tools for classification, regression, clustering, preprocessing, and evaluation.
  • Most sklearn estimators use the fit() method to learn from data.
  • Predictive models commonly use predict() to generate predictions.
  • Transformers use transform() to modify or prepare data.
  • X usually represents features and y usually represents targets.
  • Scikit-Learn works closely with NumPy and Pandas.
  • A typical workflow is prepare data → split data → preprocess → train → predict → evaluate.
  • Understanding the common sklearn API is more important than memorizing individual algorithms.
Exercise 1 · First sklearn Import

Import LinearRegression from sklearn.linear_model, print the installed sklearn version, create a LinearRegression instance, and print the model.

exercise.py
import sklearn
from sklearn.linear_model import LinearRegression

# Print sklearn version

# Create a LinearRegression instance

# Print the model

# Your code here

Challenge: After creating the model, create a small NumPy dataset, train the model using fit(), and predict the output for a new value.

Next Lecture

In Lecture 02, we will install the full data science environment, configure the required Python packages, and verify that everything works correctly before building machine learning projects.

Lecture 02 · Installation & Setup

Install and Set Up Scikit-Learn

Install Scikit-Learn and the essential Python data science libraries, create an isolated virtual environment, verify that everything works correctly, understand common installation problems, and learn the import patterns used throughout this course.

What You Will Learn
  • How to install Scikit-Learn with pip
  • How to install NumPy, Pandas, and Matplotlib
  • How to verify the installation
  • How to use virtual environments
  • Why virtual environments are important
  • Common import patterns for sklearn
  • How to check Python and pip versions
  • How to troubleshoot common installation errors
  • How to test a complete sklearn environment

01. Installing the Full Stack

Scikit-Learn works as part of the wider Python data science ecosystem. Although sklearn provides the machine learning algorithms, real machine learning projects commonly use additional libraries for numerical computation, data manipulation, and visualization.

For this course, we will primarily work with Scikit-Learn, NumPy, Pandas, and Matplotlib.

Terminal
pip install scikit-learn numpy pandas matplotlib

The command tells pip to download and install the required packages. If the packages are already installed, pip will normally report that the requirements are already satisfied or update them when appropriate.

The Four Main Libraries
  • NumPy — numerical arrays and mathematical operations
  • Pandas — data loading, cleaning, and analysis
  • Matplotlib — data visualization and charts
  • Scikit-Learn — machine learning algorithms and utilities

02. Checking Python and pip

Before installing packages, it is useful to verify that Python and pip are available from your terminal. This can help identify problems where Python is not installed correctly or where the wrong Python installation is being used.

Terminal
python --version
pip --version

You should see the installed Python version and information about the pip package manager. On some systems, the Python command may be python3 instead.

Terminal
python3 --version
pip3 --version
Tip

If your computer has multiple Python installations, always make sure that the pip you use installs packages into the same Python environment that runs your program.

03. Verifying the Installation

After installation, the next step is to verify that Python can successfully import the required libraries.

verify.py
import sklearn
import numpy as np
import pandas as pd
import matplotlib

print("sklearn:", sklearn.**version**)
print("numpy:", np.**version**)
print("pandas:", pd.**version**)
print("matplotlib:", matplotlib.**version**)

print("All installed!")

If this program runs without an import error, the basic environment is working correctly.

Why Check Versions?

Version information is useful when debugging projects. If another developer or a tutorial uses a different library version, knowing your installed versions can help explain differences in behavior.

04. Virtual Environment (Recommended)

A virtual environment creates an isolated Python environment for a project. Packages installed inside the environment are separated from packages installed in other projects.

This is particularly useful for machine learning because different projects may require different versions of libraries.

Terminal
python -m venv venv

The command creates a directory named venv containing an isolated Python environment.

On Windows, activate it using:

Windows Terminal
venv\Scripts\activate

On macOS or Linux, use:

Mac / Linux Terminal
source venv/bin/activate

Once the environment is activated, install the project dependencies inside it.

Terminal
pip install scikit-learn numpy pandas matplotlib
Why Use a Virtual Environment?
  • Keeps project dependencies isolated
  • Prevents different projects from interfering with each other
  • Makes dependency management easier
  • Helps reproduce projects on other computers
  • Reduces the risk of breaking the global Python installation

05. Activating and Deactivating the Environment

When you work on a project, activate its virtual environment before running or installing project dependencies.

On Windows:

Windows
venv\Scripts\activate

On macOS and Linux:

Mac / Linux
source venv/bin/activate

When you are finished working on the project, you can leave the virtual environment with:

Terminal
deactivate
Important

Creating a virtual environment does not automatically activate it. Make sure the correct environment is active before installing packages or running your project.

06. Standard Import Pattern

Once the environment is ready, machine learning programs typically import the specific classes and functions required for the task.

imports.py
import numpy as np
import pandas as pd

from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, accuracy_score

Notice that NumPy and Pandas use common aliases, np and pd. Scikit-Learn classes are usually imported directly from their specific modules.

Import Pattern

Prefer importing the specific sklearn classes you need instead of importing every sklearn component. This makes your code easier to read and understand.

07. Understanding sklearn Modules

Scikit-Learn is organized into different modules. Each module contains tools designed for a particular part of the machine learning workflow.

Common Modules
  • sklearn.model_selection — splitting data and selecting models
  • sklearn.preprocessing — scaling and transforming features
  • sklearn.linear_model — linear and logistic models
  • sklearn.tree — decision tree algorithms
  • sklearn.ensemble — ensemble learning algorithms
  • sklearn.neighbors — nearest-neighbor algorithms
  • sklearn.svm — support vector machine algorithms
  • sklearn.cluster — clustering algorithms
  • sklearn.metrics — evaluation metrics

Understanding this organization makes it easier to search the documentation and find the appropriate tool for a particular machine learning task.

08. Testing a Simple sklearn Model

A successful import is useful, but we can perform a stronger test by actually creating and training a model.

test_model.py
import numpy as np
from sklearn.linear_model import LinearRegression

X = np.array([
[1],
[2],
[3],
[4]
])

y = np.array([
2,
4,
6,
8
])

model = LinearRegression()

model.fit(X, y)

prediction = model.predict([[5]])

print("Prediction:", prediction[0])

If the program runs successfully and produces a prediction, your environment is capable of executing a basic Scikit-Learn machine learning workflow.

09. Installing Packages with python -m pip

On systems with multiple Python installations, using python -m pip can make it clearer which Python interpreter is being used to run pip.

Terminal
python -m pip install scikit-learn numpy pandas matplotlib

This tells Python to execute pip as a module. It is often useful when the command pip points to a different Python installation than the one you are using to run your program.

Recommended Habit

When working with virtual environments, using python -m pip can help ensure that packages are installed for the Python interpreter you intend to use.

10. Checking Installed Packages

pip can display the packages installed in the current environment.

Terminal
pip list

You can also inspect information about a particular package.

Terminal
pip show scikit-learn

The command can display information such as the installed version, package location, and dependencies.

Useful Commands
  • pip list — displays installed packages
  • pip show scikit-learn — displays information about Scikit-Learn
  • pip --version — displays pip information
  • python --version — displays the Python version

11. Common Installation Errors

Installation problems are common when setting up Python environments. Most errors can be diagnosed by checking the Python version, active environment, pip configuration, and package names.

Common Problems
  • ModuleNotFoundError — the required package may not be installed in the active environment.
  • Wrong Python environment — pip may have installed the package into a different Python installation.
  • Permission errors — the system may not allow modification of a global Python installation.
  • Network errors — pip may be unable to download packages.
  • Old pip — updating pip may resolve some installation problems.

If you receive an import error after installation, first check whether the environment where the package was installed is the same environment used to run the Python program.

12. Updating pip

Keeping pip reasonably up to date can help with package installation and dependency management.

Terminal
python -m pip install --upgrade pip

After updating pip, you can install or update your project dependencies normally.

Remember

Do not randomly upgrade every package in a production project without considering compatibility. In a real project, dependency versions should be managed deliberately.

13. Creating a Requirements File

When a project uses several Python packages, it is useful to record its dependencies in a requirements.txt file. This makes it easier to recreate the environment on another computer.

requirements.txt
numpy
pandas
matplotlib
scikit-learn

Another environment can install these dependencies using:

Terminal
pip install -r requirements.txt

Requirements files are particularly useful when sharing projects with classmates, teammates, servers, or other development machines.

14. Complete Environment Verification

We can combine the checks from this lecture into one script that verifies the important libraries and performs a simple machine learning operation.

environment_test.py
import sklearn
import numpy as np
import pandas as pd
import matplotlib

from sklearn.linear_model import LinearRegression

print("Scikit-Learn:", sklearn.**version**)
print("NumPy:", np.**version**)
print("Pandas:", pd.**version**)
print("Matplotlib:", matplotlib.**version**)

X = np.array([[1], [2], [3]])
y = np.array([2, 4, 6])

model = LinearRegression()
model.fit(X, y)

print("Test prediction:", model.predict([[4]])[0])
print("Environment is working!")

This test checks both package imports and a basic sklearn operation. If it runs without errors, your environment is ready for the upcoming machine learning lessons.

15. Best Practices for the Course

Recommended Setup
  • Create a separate virtual environment for your machine learning projects.
  • Install packages inside the active environment.
  • Use python -m pip when you need to make the Python interpreter explicit.
  • Keep a requirements.txt file for shareable projects.
  • Check package versions when troubleshooting.
  • Use meaningful project folders and keep environments isolated.
  • Do not install unnecessary packages into every project.
  • Test the environment before beginning a large machine learning project.

A clean environment saves time later. Many machine learning errors are caused not by the algorithm itself but by incorrect packages, incompatible versions, or running code in the wrong Python environment.

16. Key Points

Key Points
  • Install Scikit-Learn with pip install scikit-learn.
  • Install NumPy, Pandas, and Matplotlib for the complete basic data science stack.
  • Use python --version and pip --version to inspect your environment.
  • Virtual environments isolate project dependencies.
  • Activate the correct virtual environment before installing packages.
  • Use python -m pip when you need to ensure pip belongs to a specific Python interpreter.
  • Verify installations by importing libraries and printing their versions.
  • Scikit-Learn is organized into modules such as linear_model, preprocessing, model_selection, and metrics.
  • A simple model test can confirm that the sklearn installation actually works.
  • Requirements files make it easier to reproduce project environments.
Exercise 2 · Installation Check

Create a virtual environment named ml_env, activate it, install Scikit-Learn, NumPy, Pandas, and Matplotlib, and create a Python script that imports all four libraries.

exercise.py
import sklearn
import numpy as np
import pandas as pd
import matplotlib

print("sklearn:", sklearn.**version**)
print("numpy:", np.**version**)
print("pandas:", pd.**version**)
print("matplotlib:", matplotlib.**version**)

# Your code here

Challenge: Import LinearRegression, create a small dataset, train a model using fit(), and make one prediction. Then create a requirements.txt file containing the libraries used by your project.

Next Lecture

In Lecture 03, we will learn about NumPy arrays and the numerical data structures used by Scikit-Learn, including shapes, dimensions, indexing, and preparing data for machine learning.

Lecture 03 · The Machine Learning Workflow

Understand the Machine Learning Workflow

A machine learning model is not created by simply calling one function. Scikit-Learn provides tools that allow you to follow a complete workflow: collect or load data, prepare it, split it into training and testing sets, choose a model, train it, make predictions, evaluate its performance, and improve it when necessary.

01. Understand the Machine Learning Workflow

A typical machine learning project follows a series of steps. Each step has a specific purpose and contributes to building a reliable model.

The general workflow can be represented as:

workflow.py
1. Collect or load data
2. Explore the data
3. Prepare the data
4. Split the data
5. Choose a model
6. Train the model
7. Make predictions
8. Evaluate the model
9. Improve the model

Scikit-Learn provides functions and classes for most of these steps. For example, train_test_split() can divide the data, while classes such as LinearRegression, KNeighborsClassifier, and DecisionTreeClassifier can be used to create machine learning models.

The important idea is that training a model is only one part of the machine learning workflow. A model must also be tested on data that it did not see during training.

The Basic ML Pipeline

Data β†’ Preparation β†’ Split β†’ Model β†’ Train β†’ Predict β†’ Evaluate β†’ Improve

02. Load the Required Libraries

Before building a machine learning model, import the libraries and tools that you need. Scikit-Learn contains many machine learning algorithms and utilities.

workflow.py
import numpy as np
import pandas as pd

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score

Here, NumPy is useful for numerical operations, while Pandas is commonly used to load and work with tabular data.

From Scikit-Learn, we import train_test_split() for dividing the dataset, LinearRegression for creating a regression model, and evaluation functions for measuring how well the model performs.

You do not need to import every Scikit-Learn module in every project. Import only the tools required for the particular machine learning task.

03. Prepare Features and Targets

Machine learning models usually work with two main types of data: features and a target.

Features are the input values used by the model to make a prediction. The target is the value that the model is trying to predict.

For example, suppose we want to predict a student's exam score from the number of hours they studied.

data.py
hours = [[1], [2], [3], [4], [5], [6]]
scores = [35, 42, 50, 58, 67, 75]

Here, hours is the feature because it is used as the input to the model. The scores are the target values because they are what we want the model to predict.

Scikit-Learn normally represents features as a two-dimensional structure, even when there is only one feature. That is why the values are written as [[1], [2], [3]] rather than simply [1, 2, 3].

Remember

X usually represents the features or input data, while y usually represents the target or output that the model learns to predict.

04. Split the Data

A model should not be evaluated only on the same data that it used for learning. Instead, divide the dataset into a training set and a testing set.

The training data is used to teach the model. The testing data is kept separate and is used later to check how well the model performs on unseen examples.

Scikit-Learn provides the train_test_split() function for this purpose.

split.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
hours,
scores,
test_size=0.2,
random_state=42
)

The function returns four values. X_train and y_train contain the training data, while X_test and y_test contain the testing data.

The test_size=0.2 argument means that approximately 20% of the available data is reserved for testing, while the remaining 80% is used for training.

The random_state=42 argument makes the split reproducible. This means that running the same code again will produce the same split.

Important

Do not train your model on the test data and then use that same data to claim that the model performs well. The test set should represent data that the model has not seen during training.

05. Choose a Machine Learning Model

After preparing and splitting the data, choose an algorithm that matches the problem you are solving.

Different machine learning problems require different types of models.

Problem Example Possible Scikit-Learn Model
Regression Predict house price LinearRegression
Classification Predict spam or not spam LogisticRegression
Classification Classify flowers KNeighborsClassifier
Classification Classify customers DecisionTreeClassifier

Choosing the correct model depends on the type of problem, the dataset, the number of features, the amount of data, and the expected output.

For this example, we are predicting a numerical exam score, so regression is appropriate. We can use LinearRegression.

06. Create and Train the Model

Once a model has been selected, create an instance of the model and use the fit() method to train it.

train.py
from sklearn.linear_model import LinearRegression

model = LinearRegression()

model.fit(X_train, y_train)

The first line creates a LinearRegression object. The second line calls fit(), which allows the model to learn the relationship between the input features and the target values.

The model examines the training data and learns parameters that can later be used to make predictions.

In Scikit-Learn, the fit() method is one of the most important methods you will encounter. Many Scikit-Learn models follow the same general pattern:

fit_pattern.py
model = SomeModel()

model.fit(X_train, y_train)

Key Idea

fit() means learn from the training data. It does not mean that the model is permanently correct. The model still needs to be evaluated using data that was not used for training.

07. Make Predictions

After training, use the predict() method to generate predictions.

predict.py
predictions = model.predict(X_test)

print(predictions)

The model receives the test features stored in X_test and produces predicted values.

For example, if X_test contains the number of hours studied by students that the model has never seen before, the model will use what it learned during training to estimate their exam scores.

The predicted values are stored in the predictions variable.

Prediction can also be performed on completely new data.

new_prediction.py
new_student = [[7]]

predicted_score = model.predict(new_student)

print(predicted_score)

Here, the model receives a new input representing seven hours of study and estimates the corresponding exam score.

08. Evaluate the Model

A prediction alone does not tell us whether a model is good. We need evaluation metrics to measure its performance.

For regression problems, common metrics include Mean Squared Error (MSE) and RΒ² score.

evaluate.py
from sklearn.metrics import mean_squared_error, r2_score

mse = mean_squared_error(y_test, predictions)
r2 = r2_score(y_test, predictions)

print(β€œMSE:”, mse)
print(β€œRΒ²:”, r2)

mean_squared_error() measures the average squared difference between the actual values and predicted values. A lower MSE generally indicates that the predictions are closer to the actual values.

r2_score() measures how well the model explains the variation in the target values. A value closer to 1 generally indicates a stronger fit for a regression model.

The correct evaluation metric depends on the type of machine learning problem. Classification models use metrics such as accuracy, precision, recall, and F1-score, while regression models commonly use metrics such as MSE, MAE, and RΒ².

09. Put the Complete Workflow Together

Now we can combine the major steps into one complete Scikit-Learn program.

complete_workflow.py
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression
from sklearn.metrics import mean_squared_error, r2_score

hours = [[1], [2], [3], [4], [5], [6]]
scores = [35, 42, 50, 58, 67, 75]

X_train, X_test, y_train, y_test = train_test_split(
hours,
scores,
test_size=0.2,
random_state=42
)

model = LinearRegression()

model.fit(X_train, y_train)

predictions = model.predict(X_test)

mse = mean_squared_error(y_test, predictions)
r2 = r2_score(y_test, predictions)

print(β€œPredictions:”, predictions)
print(β€œMSE:”, mse)
print(β€œRΒ²:”, r2)

This program demonstrates the complete basic workflow. First, the data is created. Then it is divided into training and testing sets. A regression model is created and trained. The trained model makes predictions on the test data, and the predictions are finally evaluated.

Learning this workflow is more important than memorizing individual algorithms because many Scikit-Learn models use a similar interface.

10. Improve the Model

If the model does not perform well, the workflow does not simply end. Machine learning is usually an iterative process.

You may need to improve the data, select different features, preprocess the values, choose another algorithm, tune hyperparameters, or collect more training data.

A simplified improvement cycle looks like this:

improvement.py
Prepare data
      ↓
Train model
      ↓
Evaluate model
      ↓
Analyze results
      ↓
Improve data or model
      ↓
Train again
      ↓
Evaluate again

For example, a model may perform poorly because the features are not useful, the data contains missing values, the model is too simple, or the model is overfitting the training data.

Scikit-Learn provides additional tools for preprocessing, feature selection, cross-validation, and hyperparameter tuning that can be introduced as you progress through machine learning.

11. The Scikit-Learn Estimator Pattern

One of the most useful things to understand about Scikit-Learn is that many algorithms follow a consistent interface.

A typical model follows this pattern:

sklearn_pattern.py
model = Model()

model.fit(X_train, y_train)

predictions = model.predict(X_test)

This pattern makes Scikit-Learn easier to learn. Once you understand how fit() and predict() work, you can apply the same basic idea to many different algorithms.

Some models also provide methods such as score(), transform(), or fit_transform(). These methods become especially important when working with preprocessing tools and machine learning pipelines.

Machine Learning Workflow Checklist
  • Load or collect the dataset.
  • Identify the features and target.
  • Explore and understand the data.
  • Clean and prepare the data.
  • Split the data into training and testing sets.
  • Choose an appropriate machine learning algorithm.
  • Create the model.
  • Train it using fit().
  • Generate predictions using predict().
  • Evaluate the predictions with suitable metrics.
  • Improve the model when necessary.
  • Repeat the process until the model meets the required goal.
Avoid Data Leakage

Never allow information from the test set to influence the training process. This can produce unrealistically good evaluation results and make the model appear better than it really is.

As you learn preprocessing and scaling, you will see why operations such as fitting a StandardScaler should normally be learned from the training data rather than the test data.

Exercise 2 · Build an ML Workflow

Create a small machine learning program that predicts a numerical value. You can use a dataset such as study hours and exam scores, house size and house price, or advertising budget and sales.

Your program should:

  • Create or load a dataset.
  • Separate the features from the target.
  • Split the data into training and testing sets.
  • Create an appropriate Scikit-Learn model.
  • Train the model using fit().
  • Generate predictions using predict().
  • Evaluate the model using at least one suitable metric.
  • Make a prediction for a new input value.

When you finish, explain each stage of your program in your own words. You should be able to answer: What is the input? What is the target? Which data was used for training? Which data was used for testing? What did the model learn? How well did it perform?

Lecture 04 · Datasets & Data Loading

Work with Datasets in Scikit-Learn

Machine learning models need data to learn from. Before training a model, you need to know how to obtain a dataset, load it into Python, understand its structure, identify features and targets, and prepare it for the machine learning workflow. Scikit-Learn provides several built-in datasets for learning and experimentation, while Pandas makes it easy to load real-world datasets from files such as CSV.

01. What Is a Dataset?

A dataset is a collection of data that can be used to analyze a problem or train a machine learning model. A dataset usually contains multiple observations, where each observation represents one example.

For example, a student dataset might contain information about study hours, attendance, assignments, and exam scores.

students.py
Hours  Attendance  Assignments  Score
2      75          6            48
4      82          8            61
6      90          9            74
8      95          10           88

Each row represents one student, while each column represents a variable or feature.

If our goal is to predict the exam score, Hours, Attendance, and Assignments could be features, while Score would be the target.

Remember

Rows usually represent individual observations or examples, while columns represent features, variables, or the target.

02. Features and Target

Before training a machine learning model, separate the dataset into input features and the target value.

Features are the information that the model uses to make a prediction. The target is the value that the model is trying to predict.

features_target.py
import pandas as pd

data = pd.DataFrame({
β€œHours”: [2, 4, 6, 8],
β€œAttendance”: [75, 82, 90, 95],
β€œScore”: [48, 61, 74, 88]
})

X = data[[β€œHours”, β€œAttendance”]]
y = data[β€œScore”]

print(X)
print(y)

Here, X contains the input features, while y contains the target values.

Notice that X uses double square brackets because it contains multiple columns and therefore remains a two-dimensional DataFrame. The target y is a single Pandas Series.

This separation is an important step because Scikit-Learn expects the model to receive the features separately from the values it should learn to predict.

03. Use Built-in Scikit-Learn Datasets

Scikit-Learn provides several built-in datasets that are useful for learning machine learning concepts without downloading external files.

Some commonly used datasets include:

Dataset Purpose Typical Task
load_iris() Flower measurements Classification
load_digits() Handwritten digits Classification
load_breast_cancer() Medical diagnostic data Classification
load_diabetes() Diabetes-related measurements Regression
load_wine() Wine measurements Classification

These datasets are particularly useful when learning because they allow you to practice machine learning without first finding and downloading a dataset yourself.

04. Load the Iris Dataset

The Iris dataset is one of the most commonly used beginner datasets in machine learning. It contains measurements of iris flowers and information about their species.

Scikit-Learn provides it through the load_iris() function.

iris.py
from sklearn.datasets import load_iris

iris = load_iris()

print(iris)

The returned object contains several pieces of information about the dataset, including the input data, target values, feature names, and target names.

Instead of printing the entire dataset, you can access individual parts of it.

iris_data.py
from sklearn.datasets import load_iris

iris = load_iris()

print(iris.data)
print(iris.target)
print(iris.feature_names)
print(iris.target_names)

iris.data contains the input features, while iris.target contains the numerical class labels.

iris.feature_names provides the names of the features, and iris.target_names tells us which classes the numerical target values represent.

05. Understand the Dataset Shape

Before training a model, it is useful to know how much data you have and how many features are available.

NumPy arrays and Pandas DataFrames provide the shape attribute for this purpose.

shape.py
from sklearn.datasets import load_iris

iris = load_iris()

X = iris.data
y = iris.target

print(β€œFeatures shape:”, X.shape)
print(β€œTarget shape:”, y.shape)

The shape is usually displayed as a pair such as (150, 4). The first number represents the number of observations, while the second represents the number of features.

For the Iris dataset, there are 150 examples and 4 input features.

Why Check Shape?

Checking the shape helps you quickly understand your dataset and can reveal problems such as missing columns, unexpected dimensions, or incorrectly formatted input data.

06. Inspect the Data

Do not immediately train a model after loading a dataset. First inspect the data so that you understand what you are working with.

When using Pandas, common inspection methods include head(), info(), describe(), and shape.

inspect.py
import pandas as pd

data = pd.read_csv(β€œstudents.csv”)

print(data.head())
print(data.shape)
print(data.info())
print(data.describe())

head() displays the first few rows, which gives you a quick look at the data.

shape tells you how many rows and columns are present.

info() provides information about the columns, data types, and non-null values.

describe() generates statistical information for numerical columns, such as the mean, minimum, maximum, and quartiles.

These checks are useful because real-world datasets are rarely perfectly prepared for machine learning.

07. Load a CSV Dataset

Real machine learning projects often use datasets stored in files. One of the most common formats is CSV, which stands for Comma-Separated Values.

Pandas provides the read_csv() function for loading CSV files.

load_csv.py
import pandas as pd

data = pd.read_csv(β€œstudents.csv”)

print(data.head())

The string "students.csv" represents the location of the CSV file. If the file is in the same folder as your Python program, you can use just the filename.

If the file is stored somewhere else, you can provide a file path.

file_path.py
data = pd.read_csv("data/students.csv")

After loading the file, data becomes a Pandas DataFrame that can be inspected and prepared for machine learning.

08. Select Features and Target from a CSV

Once the CSV file has been loaded, select the columns that will be used as features and identify the target column.

Suppose the CSV contains the following columns:

students.csv
Hours,Attendance,Assignments,Score
2,75,6,48
4,82,8,61
6,90,9,74
8,95,10,88

If we want to predict Score, the other columns can be used as input features.

prepare.py
import pandas as pd

data = pd.read_csv(β€œstudents.csv”)

X = data[[β€œHours”, β€œAttendance”, β€œAssignments”]]
y = data[β€œScore”]

print(X)
print(y)

The variable X now contains the information used to make predictions, while y contains the values the model should learn to predict.

Choose the Target Carefully

The target should be the value you actually want the model to predict. Do not accidentally include the target column inside X, because this can cause data leakage and produce misleading results.

09. Convert Data into NumPy Arrays

Scikit-Learn works naturally with NumPy arrays and Pandas DataFrames. In many situations, you can pass a DataFrame directly to a Scikit-Learn model.

For example:

data_format.py
import pandas as pd

data = pd.read_csv(β€œstudents.csv”)

X = data[[β€œHours”, β€œAttendance”]]
y = data[β€œScore”]

print(type(X))
print(type(y))

You can also convert Pandas data into NumPy arrays when needed.

numpy_conversion.py
X_array = X.to_numpy()
y_array = y.to_numpy()

print(X_array)
print(y_array)

However, conversion is not always necessary. Most Scikit-Learn estimators can work directly with Pandas DataFrames and Series.

10. Load Different Built-in Datasets

Once you understand the Iris dataset, you can experiment with other datasets provided by Scikit-Learn.

For example, the digits dataset contains images represented as numerical features.

digits.py
from sklearn.datasets import load_digits

digits = load_digits()

X = digits.data
y = digits.target

print(β€œFeatures:”, X.shape)
print(β€œTargets:”, y.shape)

You can also load the wine dataset:

wine.py
from sklearn.datasets import load_wine

wine = load_wine()

X = wine.data
y = wine.target

print(β€œFeatures:”, X.shape)
print(β€œTargets:”, y.shape)

Experimenting with different datasets helps you understand that machine learning models can work with many types of numerical data and different prediction problems.

11. Loading Data Is Only the Beginning

Loading a dataset does not automatically make it ready for a machine learning model. Real datasets may contain missing values, duplicate rows, text values, categorical variables, extreme values, or features with very different scales.

A typical preparation process may look like this:

data_pipeline.py
Load dataset
      ↓
Inspect dataset
      ↓
Handle missing values
      ↓
Select features
      ↓
Select target
      ↓
Encode categorical data
      ↓
Scale features when necessary
      ↓
Split into training and testing data
      ↓
Train the model

Each dataset may require a different preparation process. The goal is to convert raw data into a form that a machine learning algorithm can use effectively.

Dataset Preparation Rule

Never assume that a dataset is ready just because Python successfully loaded it. Always inspect the structure, data types, missing values, and target before training a model.

12. Build a Simple Dataset-to-Model Workflow

We can now combine data loading, feature selection, splitting, training, and prediction into one simple example.

dataset_workflow.py
import pandas as pd

from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression

data = pd.read_csv(β€œstudents.csv”)

X = data[[β€œHours”, β€œAttendance”]]
y = data[β€œScore”]

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

model = LinearRegression()

model.fit(X_train, y_train)

predictions = model.predict(X_test)

print(β€œPredictions:”, predictions)

This example demonstrates an important connection between data loading and the machine learning workflow. The dataset is first loaded using Pandas, then divided into features and a target, and finally passed into the Scikit-Learn training process.

Understanding this connection will make it much easier to work with larger real-world datasets later.

Good Data Loading Habits
  • Always inspect a dataset after loading it.
  • Check the number of rows and columns.
  • Check the data types of each column.
  • Look for missing values.
  • Identify the target column before selecting features.
  • Keep features and target separate.
  • Use built-in Scikit-Learn datasets when practicing algorithms.
  • Use Pandas for convenient loading and inspection of CSV data.
  • Do not assume raw data is ready for training.
  • Prepare the data before giving it to a machine learning model.
Exercise 3 · Load and Explore a Dataset

Use one of Scikit-Learn's built-in datasets such as Iris, Digits, Wine, or Diabetes.

Your program should:

  • Import and load the dataset.
  • Display the feature names if they are available.
  • Display the target names if they are available.
  • Display the shape of the feature data.
  • Display the shape of the target data.
  • Print the first few examples.
  • Separate the features into X and the target into y.

After completing the exercise, explain what each row and column represents. Then identify whether the dataset is more suitable for a classification or regression problem.

Lecture 05 · Train / Test Split

Train / Test Split

Learn why splitting your data is essential, how training and testing data work, and how to use train_test_split correctly to evaluate machine learning models and avoid misleading results.

What You Will Learn
  • Why machine learning data must be divided into training and testing sets
  • The difference between training data and testing data
  • How to use train_test_split()
  • How to choose an appropriate test size
  • Why random_state is useful
  • How stratify preserves class proportions
  • How data splitting helps measure generalization
  • Common mistakes to avoid when splitting datasets

01. Why Split Data?

If you train and evaluate a model using exactly the same data, the model may appear to perform extremely well because it has already seen those examples. This does not tell us how well the model will perform on completely new data.

This problem is closely related to overfitting. An overfitted model learns the training examples too closely, including patterns that may not generalize to unseen data.

By separating the dataset into training and testing portions, we can train the model on one portion and evaluate it on data that was kept unseen during training.

Simple Idea

Training data teaches the model. Testing data checks whether the model learned something that generalizes to new examples.

split.py
from sklearn.model_selection import train_test_split
from sklearn.datasets import load_iris

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42,
stratify=y
)

print("Train:", X_train.shape)
print("Test:", X_test.shape)

02. Training Data vs Testing Data

The training set is the portion of the dataset used to learn the relationships between features and targets. The testing set is kept separate until evaluation.

Training Set

Used by the model to learn patterns and relationships between features and targets.

Testing Set

Contains examples that the model should not see during training.

Generalization

Describes how well a trained model performs on unseen data.

Evaluation

Measures model performance using the test data and appropriate metrics.

For example, if a dataset contains 1,000 samples and we use an 80/20 split, approximately 800 samples are used for training and 200 samples are reserved for testing.

03. Using train_test_split()

Scikit-Learn provides the train_test_split() function in sklearn.model_selection. It randomly divides one or more arrays into training and testing subsets.

basic_split.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2
)

print(X_train.shape)
print(X_test.shape)
print(y_train.shape)
print(y_test.shape)

When both X and y are supplied, Scikit-Learn keeps their corresponding rows aligned. This is extremely important because each feature row must remain associated with the correct target.

04. Understanding test_size

The test_size parameter controls how much data is placed into the testing set.

test_size.py
# 20% of the data for testing
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2
)

# 30% of the data for testing

X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3
)

# Exactly 100 samples for testing

X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=100
)
Common Values
  • test_size=0.2 — 20% testing and approximately 80% training.
  • test_size=0.25 — 25% testing and approximately 75% training.
  • test_size=0.3 — 30% testing and approximately 70% training.
  • test_size=100 — exactly 100 samples for testing.

There is no universal test size that is correct for every project. The appropriate choice depends on the dataset size, problem, and evaluation strategy.

05. Understanding random_state

By default, the split is random. Running the same program multiple times can therefore produce different training and testing samples.

The random_state parameter allows you to control the randomness. Using the same value produces the same split each time, which makes experiments easier to reproduce.

random_state.py
from sklearn.model_selection import train_test_split

X_train1, X_test1, y_train1, y_test1 = train_test_split(
X, y,
test_size=0.2,
random_state=42
)

X_train2, X_test2, y_train2, y_test2 = train_test_split(
X, y,
test_size=0.2,
random_state=42
)

print((X_train1 == X_train2).all())
Tip

The number 42 has no special machine learning meaning. It is simply a commonly used seed value. You can use another integer as long as you document it when reproducibility matters.

06. Understanding stratify

For classification problems, different classes may have different numbers of examples. A random split can sometimes produce slightly different class proportions between training and testing sets.

The stratify parameter helps preserve the class distribution when splitting the data.

stratified_split.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)

print("Training classes:", y_train)
print("Testing classes:", y_test)
Why stratify?

Suppose a classification dataset contains 70% class A and 30% class B. Stratification attempts to maintain approximately the same class proportions in both the training and testing sets.

07. Checking the Split

After splitting the data, it is good practice to check the shapes of all four resulting arrays. This confirms that the split happened as expected.

check_split.py
print("X_train:", X_train.shape)
print("X_test :", X_test.shape)
print("y_train:", y_train.shape)
print("y_test :", y_test.shape)

print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))

For the Iris dataset, there are 150 samples. With test_size=0.2, the test set contains 30 samples and the training set contains 120 samples.

08. Splitting Multiple Arrays

One important advantage of train_test_split() is that it can split multiple related arrays at the same time. This is how we keep features and targets correctly matched.

multiple_arrays.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

print("Training features:", X_train.shape)
print("Training targets:", y_train.shape)
print("Testing features:", X_test.shape)
print("Testing targets:", y_test.shape)
Important

Never independently shuffle X and y. Doing so can destroy the relationship between a feature row and its correct target. Use train_test_split(X, y, ...) so Scikit-Learn keeps them synchronized.

09. Train / Test Split with a CSV Dataset

The same technique works when the data comes from your own CSV file. First load the data using Pandas, separate the features and target, and then split them.

csv_split.py
import pandas as pd
from sklearn.model_selection import train_test_split

df = pd.read_csv("data.csv")

X = df.drop("target", axis=1)
y = df["target"]

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)

print("Training data:", X_train.shape)
print("Testing data:", X_test.shape)

10. Why We Should Not Train on Test Data

The test set should represent data that the model has never seen during training. If test data is accidentally used during training, the evaluation can become overly optimistic.

This is one form of data leakage. Data leakage occurs when information that should be unavailable to the model during training finds its way into the training process.

Avoid Data Leakage
  • Do not train the model using X_test and y_test.
  • Do not use test data to choose the best model.
  • Do not calculate preprocessing statistics from the entire dataset before splitting.
  • Keep the test set isolated until final evaluation.

11. Train / Test Split and Preprocessing

When preprocessing data, you should generally split the data first and then fit preprocessing transformations using the training data. The transformation can then be applied to the test data.

split_scale.py
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

scaler = StandardScaler()

X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
Important Pattern

Use fit_transform() on training data because the scaler must learn its parameters from the training set. Use only transform() on test data so information from the test set does not influence the learned transformation.

12. Train / Test Split vs Validation Data

For simple projects, a train/test split may be enough. However, larger machine learning workflows often divide data into three sets: training, validation, and testing.

Training Set

Used to train the model and learn its parameters.

Validation Set

Used during development to compare models and tune hyperparameters.

Testing Set

Used at the end to estimate performance on unseen data.

Cross-Validation

Provides another way to evaluate models using multiple training and validation splits.

We will study validation and cross-validation in more detail later in the course.

13. Complete Train / Test Example

The following example combines the main concepts from this lecture into a complete workflow.

complete_split.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

# Load data

X, y = load_iris(return_X_y=True)

# Split into training and testing sets

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)

# Display sizes

print("Training features:", X_train.shape)
print("Testing features:", X_test.shape)
print("Training targets:", y_train.shape)
print("Testing targets:", y_test.shape)

14. Common Mistakes

Avoid These Mistakes
  • Training and testing on the exact same data.
  • Forgetting to separate X and y.
  • Using the test set during model training.
  • Using test data to repeatedly tune the model.
  • Ignoring class imbalance in classification problems.
  • Forgetting random_state when reproducible experiments are required.
  • Fitting preprocessing transformations on the complete dataset before splitting.
  • Using an extremely small test set that does not provide a useful evaluation.

15. Key Points

Key Points
  • Training data is used to teach the model.
  • Testing data is used to evaluate the model on unseen examples.
  • train_test_split() is provided by sklearn.model_selection.
  • test_size controls the size of the test set.
  • random_state makes random splitting reproducible.
  • stratify=y helps preserve class proportions in classification tasks.
  • Never allow test data to influence the training process.
  • Data preprocessing should be fitted using training data only.
  • A proper split helps provide a more realistic estimate of model generalization.
Exercise 5 · Split a Dataset

Load the Iris dataset and split it into 80% training data and 20% testing data using random_state=0 and stratify=y. Print the shapes of all four arrays.

exercise.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)

# Split the data

# Your code here

# Print all four shapes

# X_train

# X_test

# y_train

# y_test
Next Lecture

In Lecture 06, we will learn about Data Preprocessing, including handling missing values, encoding categorical data, scaling numerical features, and preparing raw data for machine learning models.

Lecture 06 · Feature Scaling

Feature Scaling and Normalization

Learn how to scale your data so all features contribute appropriately to model training, understand why different numerical ranges can affect machine learning algorithms, and apply common Scikit-Learn scaling techniques correctly.

What You Will Learn
  • Why feature scaling is important in machine learning
  • How StandardScaler works
  • How MinMaxScaler works
  • The difference between standardization and normalization
  • Why scaling must be fitted only on training data
  • How scaling prevents data leakage
  • Which machine learning algorithms are affected by feature scale
  • How to choose an appropriate scaling technique

01. Why Feature Scaling?

Machine learning datasets often contain features measured on completely different scales. For example, one feature might represent age from 0 to 100, while another feature represents salary from 10,000 to 1,000,000.

If an algorithm depends on distances, magnitudes, or numerical optimization, a feature with much larger values can have a disproportionate influence on the model.

Simple Example

Imagine two features: age ranges from 18 to 80, while income ranges from 20,000 to 200,000. Without scaling, the numerical magnitude of income is much larger than age, even though both may be equally important.

Feature scaling transforms numerical columns to a more comparable scale while preserving the useful information contained in the data.

02. Standardization vs Normalization

Two common approaches to feature scaling are standardization and normalization. Although these terms are sometimes used interchangeably, they describe different transformations.

Standardization

Transforms features so they generally have a mean near 0 and a standard deviation near 1.

Normalization

Often refers to scaling values to a fixed range such as 0 to 1, depending on the technique being used.

StandardScaler

Uses the mean and standard deviation of the training data.

MinMaxScaler

Maps feature values to a specified range, commonly 0 to 1.

03. StandardScaler

StandardScaler is one of the most commonly used preprocessing tools in Scikit-Learn. It standardizes each feature using statistics calculated from the training data.

The transformation is based on the formula:

Standardization Formula

z = (x - mean) / standard deviation

After standardization, a feature will typically have a mean close to 0 and a standard deviation close to 1 on the data used to fit the scaler.

scaling.py
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42
)

scaler = StandardScaler()

X_train_s = scaler.fit_transform(X_train)
X_test_s  = scaler.transform(X_test)

print("Mean:", X_train_s.mean(axis=0).round(2))
print("Std: ", X_train_s.std(axis=0).round(2))

Notice that the training data uses fit_transform(), while the test data uses only transform(). This distinction is extremely important.

04. Understanding fit(), transform(), and fit_transform()

Scikit-Learn preprocessing objects follow the same API pattern used throughout the library.

scaler_methods.py
scaler = StandardScaler()

# Learn statistics from training data

scaler.fit(X_train)

# Transform the training data

X_train_s = scaler.transform(X_train)

# Transform new data using the same statistics

X_test_s = scaler.transform(X_test)
The Three Methods
  • fit() — learns parameters such as the mean and standard deviation.
  • transform() — applies the learned transformation.
  • fit_transform() — performs both operations in one call.

The common training pattern is therefore:

pattern.py
scaler.fit(X_train)

X_train_scaled = scaler.transform(X_train)
X_test_scaled = scaler.transform(X_test)

05. MinMaxScaler

MinMaxScaler transforms features into a specified range. The default range is from 0 to 1.

This can be useful when you want your features to have a predictable minimum and maximum value.

minmax.py
from sklearn.preprocessing import MinMaxScaler

scaler = MinMaxScaler()

X_train_s = scaler.fit_transform(X_train)
X_test_s  = scaler.transform(X_test)

print("Min:", X_train_s.min(axis=0))
print("Max:", X_train_s.max(axis=0))

The transformation can be represented conceptually as:

Min-Max Formula

x_scaled = (x - x_min) / (x_max - x_min)

With the default configuration, values are mapped approximately into the range 0 to 1 based on the training data.

06. Choosing a Feature Range

MinMaxScaler allows you to choose a custom output range instead of using 0 to 1.

custom_range.py
from sklearn.preprocessing import MinMaxScaler

scaler = MinMaxScaler(feature_range=(-1, 1))

X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)

print(X_train_s.min(axis=0))
print(X_train_s.max(axis=0))

This can be useful for algorithms or applications where a particular numerical range is desirable.

07. Why Scaling Helps Some Algorithms

Not every machine learning algorithm is equally affected by feature scale. Algorithms that use distances, gradients, or regularization are often particularly sensitive to differently scaled features.

Algorithms Commonly Benefiting from Scaling
  • K-Nearest Neighbors (KNN)
  • Support Vector Machines (SVM)
  • Logistic Regression
  • Linear Regression with regularization
  • Ridge and Lasso regression
  • Neural networks
  • K-Means clustering
  • Principal Component Analysis (PCA)

For example, KNN calculates distances between observations. If one feature has much larger numerical values than another, it can dominate the distance calculation.

08. Algorithms That Often Do Not Require Scaling

Some algorithms are much less sensitive to the numerical scale of individual features. Tree-based algorithms typically make decisions using thresholds rather than distance calculations.

Often Less Sensitive to Scaling
  • Decision Trees
  • Random Forests
  • Gradient-boosted tree models
  • Many other tree-based methods

This does not mean scaling is harmful in every situation. It means that scaling is generally not necessary for the core decision-making mechanism of these models.

09. Critical Rule — No Data Leakage

Critical Rule — No Data Leakage

Always fit_transform() on training data only, then transform() on the test set. Fitting on the test set causes data leakage and can give falsely optimistic evaluation results.

correct_scaling.py
scaler = StandardScaler()

# Correct

X_train_scaled = scaler.fit_transform(X_train)

# Correct

X_test_scaled = scaler.transform(X_test)

The scaler learns its mean and standard deviation from X_train. The same learned values are then used to transform X_test.

Incorrect Approach
wrong_scaling.py
scaler = StandardScaler()

X_train_scaled = scaler.fit_transform(X_train)

# Do NOT do this

X_test_scaled = scaler.fit_transform(X_test)

The second fit_transform() causes the scaler to learn new statistics from the test data. The test set should remain unseen during the fitting process.

10. Scaling After Train/Test Split

The correct order of operations is important. You should normally split the data before fitting the scaler.

correct_order.py
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

# 1. Split the data

X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42
)

# 2. Create scaler

scaler = StandardScaler()

# 3. Fit only on training data

X_train = scaler.fit_transform(X_train)

# 4. Transform test data

X_test = scaler.transform(X_test)
Correct Workflow

Load Data → Split Data → Fit Scaler on Training Data → Transform Training Data → Transform Test Data → Train Model

11. Comparing StandardScaler and MinMaxScaler

Both scalers change the numerical scale of features, but they do so differently.

StandardScaler

Centers data around zero and scales according to standard deviation.

MinMaxScaler

Maps features to a chosen range, commonly 0 to 1.

Standardization

Useful when features have different units and a zero-centered distribution is desirable.

Min-Max Scaling

Useful when a bounded numerical range is important for the application.

12. Inspecting the Scaler

After fitting a scaler, Scikit-Learn stores the learned parameters inside the scaler object. These values can be inspected to understand what the preprocessing step learned from the training data.

inspect_scaler.py
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()

X_train_scaled = scaler.fit_transform(X_train)

print("Means:", scaler.mean_)
print("Scales:", scaler.scale_)

These values are learned from the training data and are then reused whenever new data is transformed using the same scaler.

13. Scaling a New Sample

Once a scaler has been fitted, the same scaler can be used to transform new observations before sending them to a trained model.

new_data.py
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()

X_train_scaled = scaler.fit_transform(X_train)

new_sample = X_test[:1]

new_sample_scaled = scaler.transform(new_sample)

print("Original:", new_sample)
print("Scaled:", new_sample_scaled)
Important

Do not create and fit a brand-new scaler for every new sample. Reuse the scaler that was fitted during model training so the new data is transformed using the same rules.

14. Complete Scaling Example

The following example combines dataset loading, train/test splitting, scaling, and a simple model into one workflow.

complete_scaling.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

X, y = load_iris(return_X_y=True)

X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42,
stratify=y
)

scaler = StandardScaler()

X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

model = KNeighborsClassifier(n_neighbors=3)

model.fit(X_train, y_train)

predictions = model.predict(X_test)

print("Accuracy:", accuracy_score(y_test, predictions))

15. Common Scaling Mistakes

Avoid These Mistakes
  • Fitting the scaler on the entire dataset before splitting.
  • Calling fit_transform() on the test data.
  • Creating a different scaler for training and testing data.
  • Forgetting to scale new prediction data using the same fitted scaler.
  • Assuming every algorithm requires feature scaling.
  • Ignoring data leakage during preprocessing.
  • Scaling categorical values without first choosing an appropriate encoding strategy.

16. Key Points

Key Points
  • Feature scaling puts numerical features on a more comparable scale.
  • StandardScaler standardizes features using the mean and standard deviation.
  • MinMaxScaler maps features into a specified range.
  • Distance-based and gradient-based algorithms often benefit from scaling.
  • Many tree-based algorithms are less sensitive to feature scale.
  • Always fit preprocessing objects using training data only.
  • Use fit_transform() on training data.
  • Use transform() on test and future data.
  • Incorrect scaling can cause data leakage.
  • The same fitted scaler should be reused when transforming new observations.
Exercise 6 · Scale the Iris Dataset

Load Iris, split the data into 80% training and 20% testing sets, apply StandardScaler correctly, and print the mean and standard deviation of the scaled training set.

exercise.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler

X, y = load_iris(return_X_y=True)

# Split the data

# Your code here

# Create the scaler

# Your code here

# Scale training and testing data

# Your code here

# Print the mean and standard deviation

# Your code here
Next Lecture

In Lecture 07, we will learn about Encoding Categorical Data, including how to convert text-based categories into numerical representations that machine learning models can understand.

Lecture 07 · 🛠️ Project 1

Project 1: Iris Flower Classifier

Apply everything from Unit I to build your first complete classification model with Scikit-Learn. You will load real data, explore it, split it correctly, scale the features, train a KNN model, make predictions, and evaluate its performance.

Project Goals
  • Load and explore the Iris dataset
  • Understand features and target labels
  • Split the dataset into training and testing sets
  • Scale numerical features with StandardScaler
  • Understand how K-Nearest Neighbors works
  • Train a KNN classification model
  • Make predictions on unseen test data
  • Evaluate the model using accuracy
  • Generate a classification report
  • Test the trained model on new flower measurements

01. Project Overview

In this project, we will build a machine learning model that can identify the species of an Iris flower from measurements of its physical characteristics.

The Iris dataset contains measurements of iris flowers belonging to three species: setosa, versicolor, and virginica.

Sepal Length

Length of the flower's sepal measured in centimeters.

Sepal Width

Width of the flower's sepal measured in centimeters.

Petal Length

Length of the flower's petal measured in centimeters.

Petal Width

Width of the flower's petal measured in centimeters.

These four measurements become our features, while the flower species becomes our target.

02. The Machine Learning Pipeline

This project combines the main concepts covered in Unit I into one complete workflow.

Our Pipeline

Load Data → Explore Data → Split Data → Scale Features → Create Model → Train Model → Predict → Evaluate

Each step has a specific purpose. Keeping these steps separate makes machine learning projects easier to understand, debug, and maintain.

03. Loading the Iris Dataset

Scikit-Learn provides the Iris dataset through load_iris(). We can load it directly without downloading a CSV file.

load_iris.py
from sklearn.datasets import load_iris

iris = load_iris()

print("Feature names:", iris.feature_names)
print("Target names:", iris.target_names)
print("Data shape:", iris.data.shape)
print("Target shape:", iris.target.shape)

The dataset contains 150 samples and 4 numerical features. The target contains three possible flower classes.

04. Separating X and y

Machine learning models generally use X to represent input features and y to represent target values.

features.py
iris = load_iris()

X = iris.data
y = iris.target

print("X shape:", X.shape)
print("y shape:", y.shape)

print("First X row:", X[0])
print("First y value:", y[0])
Understanding X and y
  • X contains the four flower measurements.
  • y contains the numerical class labels.
  • X.shape is (150, 4).
  • y.shape is (150,).

05. Splitting the Dataset

We should not train and evaluate our model using exactly the same examples. Instead, we reserve part of the dataset for testing.

train_test.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)

print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))

With 150 total samples and a 20% test size, approximately 120 samples are used for training and 30 samples are reserved for testing.

Why stratify=y?

stratify=y helps maintain the class proportions when creating the training and testing sets. This is especially useful for classification problems.

06. Scaling the Features

KNN is a distance-based algorithm. This means the numerical scale of the features can affect how distances between flowers are calculated.

We therefore use StandardScaler to standardize the features.

scaling.py
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler()

X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

print("Scaled training data:")
print(X_train[:5])
Remember the Scaling Rule

Use fit_transform() on the training data and only transform() on the test data. Never fit the scaler separately on the test set.

07. Understanding K-Nearest Neighbors

K-Nearest Neighbors (KNN) is a supervised machine learning algorithm that classifies a new observation based on nearby training observations.

When a new flower is given to the model, KNN looks at the closest training examples and uses their classes to determine the predicted class.

Simple KNN Example

If most of the five nearest flowers belong to the versicolor class, the model will generally predict the new flower as versicolor.

08. Choosing n_neighbors

The n_neighbors parameter determines how many nearby training samples KNN considers when making a prediction.

knn_values.py
from sklearn.neighbors import KNeighborsClassifier

model_3 = KNeighborsClassifier(n_neighbors=3)
model_5 = KNeighborsClassifier(n_neighbors=5)
model_10 = KNeighborsClassifier(n_neighbors=10)

Using a very small value can make the model highly sensitive to individual training examples, while a larger value considers a broader neighborhood.

For This Project

We will use n_neighbors=5. Later in the course, you will learn how to systematically choose model hyperparameters instead of selecting them manually.

09. Creating and Training the Model

Once our data has been prepared, we create a KNN classifier and train it using the training data.

train_model.py
from sklearn.neighbors import KNeighborsClassifier

model = KNeighborsClassifier(n_neighbors=5)

model.fit(X_train, y_train)

print("Model trained successfully!")

The fit() method gives the model access to the training features and their corresponding target labels.

10. Making Predictions

After training, we can use the model to predict the classes of flowers in the test set.

predict.py
y_pred = model.predict(X_test)

print("Predicted:", y_pred)
print("Actual:   ", y_test)

The predicted values are compared with the actual target values to determine how well the model performed.

11. Measuring Accuracy

Accuracy measures the proportion of predictions that were correct.

accuracy.py
from sklearn.metrics import accuracy_score

accuracy = accuracy_score(y_test, y_pred)

print("Accuracy:", accuracy)
print("Accuracy percentage:", accuracy * 100, "%")
Understanding Accuracy

If the model correctly predicts 29 out of 30 test samples, its accuracy is 29 / 30 = approximately 96.7%.

12. Classification Report

Accuracy gives us one overall measurement, but classification problems often require more detailed evaluation. Scikit-Learn provides classification_report() for this purpose.

report.py
from sklearn.metrics import classification_report

print(
classification_report(
y_test,
y_pred,
target_names=iris.target_names
)
)

The report includes metrics such as precision, recall, and F1-score for each class.

Important Metrics
  • Precision — how many predicted positives were actually correct.
  • Recall — how many actual positives were successfully identified.
  • F1-score — combines precision and recall into a single metric.
  • Support — number of actual samples belonging to each class.

13. Complete Project Code

Now we can combine everything into one complete Python program.

iris_classifier.py
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score, classification_report

# Load dataset

iris = load_iris()

X, y = iris.data, iris.target

print("Shape:", X.shape)
print("Classes:", iris.target_names)

# Split data

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)

# Scale features

scaler = StandardScaler()

X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)

# Create model

model = KNeighborsClassifier(n_neighbors=5)

# Train model

model.fit(X_train, y_train)

# Predict

y_pred = model.predict(X_test)

# Evaluate

accuracy = accuracy_score(y_test, y_pred)

print("Accuracy:", accuracy)

print(classification_report(
y_test,
y_pred,
target_names=iris.target_names
))

14. Making a Prediction for a New Flower

One of the most useful parts of machine learning is making predictions on new data. Suppose we receive measurements for a flower that was not part of the original training dataset.

new_flower.py
new_flower = [[5.1, 3.5, 1.4, 0.2]]

new_flower_scaled = scaler.transform(new_flower)

prediction = model.predict(new_flower_scaled)

print("Predicted class:", prediction[0])
print("Predicted species:", iris.target_names[prediction[0]])
Important

The new flower must be transformed using the same scaler that was fitted on the training data. Do not create and fit a new scaler for the new sample.

15. Understanding the Complete Pipeline

At this point, the entire project can be understood as a sequence of connected steps.

1. Load

Load the Iris dataset using Scikit-Learn.

2. Prepare

Separate features X and target y.

3. Split

Create training and testing datasets.

4. Scale

Standardize numerical features.

5. Train

Fit the KNN classifier on training data.

6. Predict

Predict classes for unseen test samples.

7. Evaluate

Measure accuracy and classification metrics.

8. Deploy Logic

Use the trained model and scaler to classify new flowers.

16. Expected Result

Expected Result

You should typically see approximately 96–100% accuracy on the Iris dataset with KNN using this particular split and configuration. Exact results can vary when the train/test split, random seed, preprocessing, or model parameters change.

A high accuracy on this small, well-known dataset does not mean that every real-world classification problem will be equally easy. Real datasets can contain missing values, noise, class imbalance, outliers, irrelevant features, and much more complex relationships.

17. Common Mistakes

Avoid These Mistakes
  • Training the model before splitting the dataset.
  • Fitting the scaler on the entire dataset.
  • Using fit_transform() on the test data.
  • Forgetting to scale new observations before prediction.
  • Mixing up X and y.
  • Evaluating predictions against the wrong target values.
  • Assuming high accuracy on Iris means every classification problem will be easy.
  • Changing model parameters repeatedly based on the test set without a proper validation strategy.

18. Key Points

Key Points
  • The Iris dataset contains 150 samples, 4 features, and 3 classes.
  • X contains the flower measurements and y contains the target classes.
  • train_test_split() separates training and testing data.
  • StandardScaler standardizes the numerical features.
  • KNN classifies samples based on nearby training observations.
  • n_neighbors controls how many neighbors KNN considers.
  • fit() trains the KNN model.
  • predict() generates predictions for new observations.
  • accuracy_score() measures overall classification accuracy.
  • classification_report() provides detailed classification metrics.
  • The same fitted scaler must be used when processing future data.
Exercise 7 · Improve the Iris Classifier

Modify the project by testing different values of n_neighbors. Try 3, 5, 7, and 9. Train each model using the same training data and compare their test accuracy.

exercise.py
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score

for k in [3, 5, 7, 9]:

```
model = KNeighborsClassifier(n_neighbors=k)

# Train the model
model.fit(X_train, y_train)

# Make predictions
y_pred = model.predict(X_test)

# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)

print("K =", k, "| Accuracy =", accuracy)
```
Next Lecture

In Lecture 08, we will begin the next stage of the machine learning workflow: Model Evaluation. You will learn how to interpret evaluation metrics, confusion matrices, precision, recall, F1-score, and understand whether a model is actually performing well.

Lecture 08 · Linear Regression

Linear Regression

Learn how to predict continuous values using Linear Regression, the most fundamental supervised learning algorithm. You will understand the mathematical idea behind a regression line, train a model with scikit-learn, evaluate its performance, and interpret what the model has learned.

What You Will Learn
  • What linear regression is and how it works
  • The difference between regression and classification
  • How a regression line makes predictions
  • How to train LinearRegression in sklearn
  • How coefficients and intercepts are calculated and interpreted
  • How to make predictions for new data
  • How to evaluate regression models with MSE and R²
  • How to understand underfitting and poor model performance
  • How to avoid common regression mistakes

01. What Is Linear Regression?

Linear Regression is a supervised machine learning algorithm used to predict a continuous numerical value.

For example, we can use Linear Regression to predict:

  • House price from house size
  • Salary from years of experience
  • Temperature from time of day
  • Sales from advertising expenditure
  • Student marks from study hours
  • Fuel consumption from vehicle characteristics

The word linear means that the model attempts to describe the relationship between input variables and the target using a straight-line relationship.

Regression vs Classification

Regression predicts numerical values such as 250000, 72.5, or 31.8.

Classification predicts categories such as "cat", "dog", "spam", or "not spam".

Our Iris project used classification because the target was a flower species. Linear Regression is different because its output is a continuous number.

02. The Basic Linear Regression Equation

For a single input feature, Linear Regression uses the equation:

y = mx + b

  • y = predicted value
  • x = input feature
  • m = slope or coefficient
  • b = intercept

In machine learning, this can also be written as:

Ε· = Ξ²β‚€ + β₁x

Here, Ε· represents the predicted value, Ξ²β‚€ is the intercept, and β₁ is the coefficient of the feature.

equation_example.py
x = 5
m = 10
b = 5

y = m * x + b

print(y)

The result is 55 because:

y = (10 Γ— 5) + 5 = 55

The trained Linear Regression model automatically learns suitable values for the coefficient and intercept from the training data.

03. Understanding the Coefficient

The coefficient tells us how much the predicted target changes when the input feature increases by one unit.

Suppose our model learns:

y = 8x + 20

The coefficient is 8.

This means that for every one-unit increase in x, the predicted value of y increases by approximately 8 units.

Important

A positive coefficient means the prediction generally increases as the feature increases. A negative coefficient means the prediction generally decreases as the feature increases. A coefficient close to zero indicates a weak linear relationship, although its practical meaning depends on the scale and context of the feature.

04. Understanding the Intercept

The intercept is the predicted value of the target when all input features are zero.

For example:

y = 10x + 15

The intercept is 15.

When x = 0:

y = (10 Γ— 0) + 15 = 15

In real-world datasets, the intercept may not always have a meaningful practical interpretation if a value of zero for the feature is outside the realistic range of the data.

05. How Linear Regression Learns

During training, Linear Regression searches for coefficients that produce predictions as close as possible to the actual target values.

The difference between an actual value and a predicted value is called the residual or error.

The model commonly uses the ordinary least squares approach, which finds parameters that minimize the sum of squared residuals.

Conceptually:

Residual = Actual Value βˆ’ Predicted Value

Squaring the errors prevents positive and negative errors from cancelling each other out and gives larger errors more influence during optimization.

Goal of Training

The model tries to find the line that provides the smallest overall squared prediction error for the training data.

06. Creating Training Data

Before training a regression model, we need two important pieces of information:

  • X β€” the input features
  • y β€” the target values we want to predict

For example, imagine that we want to predict a person's score based on the number of hours they studied.

data.py
import numpy as np

X = np.array([[1], [2], [3], [4], [5]])
y = np.array([20, 35, 45, 55, 70])

print(X)
print(y)

Here, X contains the number of study hours, while y contains the corresponding scores.

Shape Matters

Scikit-learn expects the feature matrix X to normally have two dimensions: (number of samples, number of features). That is why we use [[1], [2], [3]] instead of [1, 2, 3] for X.

07. Importing LinearRegression

Scikit-learn provides Linear Regression through the sklearn.linear_model module.

import.py
from sklearn.linear_model import LinearRegression

We can then create a Linear Regression model:

model.py
model = LinearRegression()

At this point, the model has been created but it has not learned anything yet. Learning happens when we call fit().

08. Splitting the Dataset

Just like the Iris classification project, we should separate our data into training and testing sets.

split.py
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

The training data is used to learn the relationship between X and y. The testing data is kept separate so that we can evaluate how well the model performs on unseen examples.

Why Keep Test Data Separate?

If we evaluate the model using the same data it used for learning, we may get an overly optimistic view of its performance. A separate test set gives us a better estimate of how the model behaves on new data.

09. Training the Model

Once the data has been split, we can train the Linear Regression model using fit().

train.py
model = LinearRegression()

model.fit(X_train, y_train)

During fit(), scikit-learn calculates the model parameters based on the training data.

After training, we can inspect the values learned by the model:

parameters.py
print("Coefficient:", model.coef_)
print("Intercept:", model.intercept_)

10. Making Predictions

Once the model has been trained, we can use predict() to generate predictions for new data.

predict.py
y_pred = model.predict(X_test)

print(y_pred)

We can also predict the value for a completely new input.

new_prediction.py
new_data = np.array([[11]])

prediction = model.predict(new_data)

print("Prediction:", prediction[0])

The model uses the equation it learned from the training data to calculate the predicted value.

11. Mean Squared Error (MSE)

Mean Squared Error, commonly called MSE, measures the average squared difference between actual values and predicted values.

The formula is:

MSE = (1/n) Γ— Ξ£(actual βˆ’ predicted)Β²

A lower MSE generally indicates that the predictions are closer to the actual values.

mse.py
from sklearn.metrics import mean_squared_error

mse = mean_squared_error(y_test, y_pred)

print("MSE:", mse)
Remember

MSE is expressed in squared units of the target. Therefore, its numerical value can sometimes be difficult to interpret directly. The scale of the target variable matters when deciding whether an MSE is large or small.

12. R² Score

R², or the coefficient of determination, measures how much of the variation in the target is explained by the regression model relative to a baseline that always predicts the mean target value.

A commonly used interpretation is:

  • R² = 1.0 — perfect predictions on the evaluated data
  • R² = 0 — equivalent to the mean-prediction baseline
  • R² < 0 — worse than that baseline on the evaluated data
r2.py
from sklearn.metrics import r2_score

score = r2_score(y_test, y_pred)

print("R2 Score:", score)
Important

R² should not be interpreted as a percentage of predictions that are correct. It describes explained variance relative to a mean-based baseline and must be interpreted in the context of the dataset and model.

13. Complete Linear Regression Example

Now let's put the complete workflow together: create data, split it, train the model, make predictions, and evaluate the results.

linear_regression.py
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score

X = np.array([[1],[2],[3],[4],[5],[6],[7],[8],[9],[10]])
y = np.array([15, 25, 35, 45, 55, 62, 70, 79, 88, 95])

X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)

model = LinearRegression()

model.fit(X_train, y_train)

y_pred = model.predict(X_test)

print("Coefficient:", model.coef_[0])
print("Intercept:", model.intercept_)

print("Predictions:", y_pred)
print("Actual:", y_test)

print("MSE:", mean_squared_error(y_test, y_pred))
print("R2 Score:", r2_score(y_test, y_pred))

This example follows the standard supervised learning workflow:

Data β†’ Split β†’ Train β†’ Predict β†’ Evaluate

14. Interpreting the Model

Suppose the model produces:

output.txt
Coefficient: 8.9
Intercept: 6.2
MSE: 4.5
R2 Score: 0.97

The coefficient tells us that an increase of one unit in X is associated with an increase of approximately 8.9 units in the predicted target, according to this fitted model.

The intercept is approximately 6.2, meaning the model predicts 6.2 when X is zero.

The R² value of 0.97 indicates that the fitted model explains a large proportion of the target variation relative to the mean baseline on this evaluated dataset. It does not mean that every prediction is exactly 97% correct.

15. Multiple Linear Regression

Linear Regression is not limited to one feature. We can use multiple features to predict a single continuous target.

For example, house price could depend on:

  • House size
  • Number of bedrooms
  • Number of bathrooms
  • Age of the house

The equation becomes:

y = Ξ²β‚€ + β₁x₁ + Ξ²β‚‚xβ‚‚ + β₃x₃ + ... + Ξ²β‚™xβ‚™

multiple_regression.py
import numpy as np
from sklearn.linear_model import LinearRegression

X = np.array([
[1000, 2],
[1500, 3],
[2000, 3],
[2500, 4],
[3000, 5]
])

y = np.array([150000, 220000, 300000, 380000, 450000])

model = LinearRegression()
model.fit(X, y)

print("Coefficients:", model.coef_)
print("Intercept:", model.intercept_)

new_house = np.array([[2200, 3]])

prediction = model.predict(new_house)

print("Predicted price:", prediction[0])

Here, the first feature represents house size and the second feature represents the number of bedrooms. The model learns a separate coefficient for each feature.

16. Visualizing a Regression Line

For a dataset with one feature, visualization is an excellent way to understand Linear Regression. We can plot the original data points and then draw the line learned by the model.

plot_regression.py
import matplotlib.pyplot as plt

plt.scatter(X, y, label="Actual Data")
plt.plot(X, model.predict(X), label="Regression Line")

plt.xlabel("X")
plt.ylabel("y")
plt.title("Linear Regression")

plt.legend()
plt.show()

The dots represent the observed training examples, while the line represents the predictions made by the model.

Visual Intuition

If the data points roughly follow a straight-line pattern, Linear Regression may be a useful starting model. If the relationship is strongly curved or much more complex, a simple linear model may not capture the pattern adequately.

17. Common Linear Regression Mistakes

Beginners often make several mistakes when building regression models.

  • Using classification metrics: Accuracy is generally not the appropriate primary metric for continuous regression targets.
  • Testing on training data: This can give an overly optimistic evaluation.
  • Ignoring data leakage: Information from the test set should not influence training or preprocessing decisions.
  • Assuming high R² means a perfect model: R² must be interpreted together with the prediction errors and application context.
  • Ignoring outliers: Extreme observations can strongly influence ordinary least squares regression.
  • Assuming correlation proves causation: A fitted relationship does not by itself establish that one variable causes another.
  • Using linear regression for every problem: Some relationships are not well represented by a straight-line model.
Avoid Data Leakage

Always separate the training and testing data before making data-driven preprocessing decisions. The test set should behave like unseen future data during model development.

18. Key Points to Remember

Regression

Predicts continuous numerical values rather than discrete classes.

Coefficient

Describes how the prediction changes with a one-unit change in a feature, holding other features constant in a multiple regression model.

Intercept

The model's predicted target when all input features are zero.

MSE

Measures the average squared prediction error. Lower values are generally better.

R²

Measures performance relative to a mean-based baseline.

fit()

Trains the Linear Regression model using the training data.

predict()

Uses the learned model to generate predictions for new feature values.

Test Data

Provides an independent evaluation of how the model performs on unseen examples.

Evaluation Metrics for Regression
  • MSE — average squared difference between actual and predicted values. Lower is better when comparing models on the same target scale.
  • R² Score — measures performance relative to predicting the mean target value. 1.0 represents a perfect fit on the evaluated data, while negative values are possible.
Exercise 8 · House Size to Price

Create a dataset where X = house sizes (500–3000 sqft) and y = prices.

Train a LinearRegression model and perform the following tasks:

  • Split the dataset into training and testing sets.
  • Train the Linear Regression model.
  • Print the coefficient and intercept.
  • Predict the price of a 2000 sqft house.
  • Calculate the MSE.
  • Calculate the R² score.
  • Try predicting prices for 1000, 1500, 2500, and 3000 sqft houses.
  • Explain what the coefficient means in the context of house prices.
Challenge

Use Matplotlib to create a scatter plot of the house sizes and prices. Then draw the regression line on the same graph. Look at the points and decide whether a straight-line relationship appears reasonable.

Next Lecture

In the next lecture, we will explore another important supervised learning algorithm: Logistic Regression. You will learn how regression can also be used for classification problems by predicting probabilities and class labels.