Introduction to Scikit-Learn
Learn what Scikit-Learn is, why it is one of the most widely used machine learning libraries in Python, how it fits into the Python data science ecosystem, and how its simple API makes machine learning easier to build and understand.
- What Scikit-Learn (sklearn) is
- Why sklearn is widely used in machine learning
- What problems sklearn solves
- How to install sklearn
- How to import sklearn modules
- The sklearn API pattern: fit / predict / transform
- The difference between features and labels
- What estimators, transformers, and predictors are
- How sklearn works with NumPy and Pandas
- How to create your first machine learning model
01. What is Scikit-Learn?
Scikit-Learn (imported as sklearn) is a powerful open-source machine learning library for Python. It provides simple and efficient tools for data mining, data analysis, data preprocessing, machine learning, model evaluation, and building predictive systems.
Scikit-Learn is designed to make common machine learning tasks easier. Instead of implementing algorithms such as linear regression, logistic regression, decision trees, support vector machines, or k-nearest neighbors from scratch, developers can use tested implementations provided by the library.
It is built on top of important Python scientific computing libraries such as NumPy and SciPy, and it works extremely well with Pandas for handling datasets and Matplotlib for visualization.
Scikit-Learn is a Python machine learning library that provides ready-to-use tools for preparing data, training models, making predictions, and evaluating machine learning systems.
02. Why Use Scikit-Learn?
Machine learning involves much more than simply choosing an algorithm. A complete machine learning workflow usually includes loading data, cleaning data, preprocessing features, splitting datasets, training models, evaluating results, tuning parameters, and making predictions.
Scikit-Learn provides tools for almost every traditional machine learning step in a consistent and easy-to-learn API.
- One consistent API for many machine learning algorithms
- Comprehensive tools for classification, regression, clustering, and dimensionality reduction
- Powerful preprocessing and feature transformation utilities
- Model evaluation and cross-validation tools
- Hyperparameter tuning capabilities
- Free and open source
- Excellent documentation and a large community
- Works seamlessly with NumPy, Pandas, SciPy, and Matplotlib
- Useful for both learning machine learning and developing real applications
03. What Can Scikit-Learn Do?
Scikit-Learn covers many of the most common tasks encountered in traditional machine learning. The library contains implementations of supervised and unsupervised learning algorithms as well as tools for preparing and evaluating data.
Classification
Predict which category something belongs to. Examples include spam detection, customer churn prediction, sentiment classification, and image classification.
Regression
Predict a continuous numerical value. Examples include house price prediction, sales prediction, temperature prediction, and demand forecasting.
Clustering
Group similar data points without predefined labels. Examples include customer segmentation, document grouping, and pattern discovery.
Preprocessing
Scale numerical features, encode categorical values, handle transformations, and prepare raw data for machine learning algorithms.
Model Evaluation
Measure model performance using accuracy, precision, recall, F1-score, mean squared error, cross-validation, confusion matrices, and other metrics.
Dimensionality Reduction
Reduce the number of input features while attempting to preserve useful information using techniques such as PCA.
04. Installing Scikit-Learn
Scikit-Learn can be installed using Python's package manager, pip. It is normally installed inside the Python environment where you plan to run your machine learning programs.
pip install scikit-learn numpy pandas matplotlib
The package is installed using the name scikit-learn, but it is imported in Python using the name sklearn.
Installation name: scikit-learn
Python import name: sklearn
If you are working inside a virtual environment, activate that environment before installing the package. This helps keep project dependencies isolated from other Python projects.
05. Importing Scikit-Learn
You normally do not need to import the entire Scikit-Learn library manually. Instead, import the specific class or function required for your machine learning task.
from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler import sklearn print(sklearn.**version**)
This approach keeps your code organized and makes it clear which machine learning components your program is using.
Scikit-Learn is divided into different modules. For example, sklearn.linear_model contains linear models, sklearn.preprocessing contains preprocessing tools, and sklearn.model_selection contains tools for splitting data and performing model selection.
06. The sklearn API Pattern
One of the greatest strengths of Scikit-Learn is its consistent API. Many algorithms use similar methods even though they solve completely different machine learning problems.
The most important methods you will encounter are fit(), predict(), and transform().
from sklearn.linear_model import LinearRegression # Step 1: Create the model model = LinearRegression() # Step 2: Train the model on training data model.fit(X_train, y_train) # Step 3: Make predictions on new data predictions = model.predict(X_test) print(predictions)
fit(X, y)— learns patterns or parameters from training datapredict(X)— uses a trained model to generate predictionstransform(X)— changes or prepares data using a transformer
The fit() method is particularly important because this is where the machine learning algorithm learns from the training data. After fitting, the model can usually be used with predict() to make predictions on previously unseen data.
07. Features and Labels
Machine learning models usually work with two important types of information: features and labels.
Features are the input variables used by a model. A label, also called a target, is the value the model is trying to predict.
import numpy as np # Features: house size in square feet X = np.array([ [1000], [1500], [2000], [2500] ]) # Labels: house prices y = np.array([ 150000, 220000, 300000, 360000 ])
Here, X contains the input feature, which is the size of the house. The array y contains the target values, which are the corresponding house prices.
- X usually represents input features.
- y usually represents the target or label.
- The model learns a relationship between X and y during training.
- New values of X can then be supplied to generate predictions.
08. Estimators, Transformers, and Predictors
Scikit-Learn uses a few important concepts to organize its machine learning tools.
- Estimator — an sklearn object that learns parameters from data using
fit(). - Predictor — an estimator that can make predictions using
predict(). - Transformer — an object that modifies data using methods such as
transform(). - Model — commonly refers to a trained machine learning estimator.
For example, LinearRegression is an estimator and predictor because it can learn from data and then make numerical predictions. StandardScaler is a transformer because it learns scaling parameters and uses them to transform features.
09. A Complete First sklearn Example
Let's combine the basic concepts and build a small linear regression model. The example uses house size as the feature and house price as the target.
import numpy as np
from sklearn.linear_model import LinearRegression
# Training data
X = np.array([
[1000],
[1500],
[2000],
[2500]
])
y = np.array([
150000,
220000,
300000,
360000
])
# Create the model
model = LinearRegression()
# Train the model
model.fit(X, y)
# Predict the price of a 1800 square foot house
prediction = model.predict([[1800]])
print("Predicted price:", prediction[0])This example demonstrates the fundamental sklearn workflow: prepare data, create an estimator, call fit(), and then call predict().
Do not focus on memorizing every machine learning algorithm at the beginning. First understand the common sklearn workflow. Once you understand fit(), predict(), preprocessing, and evaluation, learning individual algorithms becomes much easier.
10. Scikit-Learn and NumPy
Scikit-Learn works closely with NumPy. Machine learning datasets are commonly represented using NumPy arrays or similar array-like structures.
import numpy as np from sklearn.linear_model import LinearRegression X = np.array([ [1], [2], [3], [4] ]) y = np.array([ 2, 4, 6, 8 ]) model = LinearRegression() model.fit(X, y) print(model.predict([[5]]))
NumPy provides the numerical array structure, while Scikit-Learn provides the machine learning algorithms and workflow.
11. Scikit-Learn and Pandas
Pandas is commonly used to load, clean, inspect, and prepare datasets before sending them to Scikit-Learn.
import pandas as pd
from sklearn.linear_model import LinearRegression
data = pd.DataFrame({
"hours": [1, 2, 3, 4, 5],
"scores": [45, 50, 60, 70, 80]
})
X = data[["hours"]]
y = data["scores"]
model = LinearRegression()
model.fit(X, y)
print(model.predict([[6]]))This is a common real-world workflow. Pandas manages the dataset, while Scikit-Learn handles the machine learning model.
Pandas → load and clean data → NumPy / Pandas → prepare features → Scikit-Learn → train and evaluate model → Matplotlib → visualize results.
12. Training and Testing Data
A machine learning model should not normally be evaluated only on the same data used to train it. Instead, a dataset is commonly divided into training and testing portions.
The training data is used to learn patterns. The testing data is kept separate and used to estimate how well the trained model performs on unseen examples.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)
print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))Never evaluate a model only on the same examples it was trained on when you want to understand how it generalizes to unseen data. Separate evaluation data helps provide a more realistic measurement of model performance.
13. The Machine Learning Workflow
Scikit-Learn is most useful when it is understood as part of a complete machine learning workflow rather than simply a collection of algorithms.
Collect Data
β
Explore Data
β
Clean Data
β
Prepare Features
β
Split Training / Testing Data
β
Preprocess Data
β
Create Model
β
Train with fit()
β
Predict with predict()
β
Evaluate Model
β
Improve and Tune ModelEach stage is important. A powerful algorithm cannot compensate for badly prepared data, data leakage, incorrect evaluation, or poorly selected features.
14. Common Scikit-Learn Module Categories
Scikit-Learn contains many modules. Learning how these modules are organized makes it easier to find the right tool.
sklearn.linear_model— linear and logistic regression modelssklearn.tree— decision tree algorithmssklearn.ensemble— ensemble algorithms such as random forestssklearn.neighbors— nearest-neighbor algorithmssklearn.svm— support vector machine algorithmssklearn.cluster— clustering algorithmssklearn.preprocessing— scaling, encoding, and feature transformationssklearn.model_selection— train/test splitting and model selectionsklearn.metrics— model evaluation metricssklearn.decomposition— dimensionality reduction techniques
15. Common Beginner Mistakes
Beginners often make mistakes when learning machine learning libraries. Understanding these mistakes early will make your sklearn programs more reliable.
- Forgetting to install Scikit-Learn in the correct Python environment
- Using
scikit-learninstead ofsklearnin an import statement - Training and evaluating a model on exactly the same data
- Forgetting to preprocess features when the algorithm requires it
- Ignoring missing values or invalid data
- Using inappropriate evaluation metrics
- Assuming a more complicated model is automatically better
- Ignoring data leakage during preprocessing and evaluation
Machine learning is not simply about obtaining a prediction. You must also understand how the data was prepared, how the model was trained, and how its performance was measured.
16. Key Points
- Scikit-Learn is a major Python library for traditional machine learning.
- Install it using
pip install scikit-learn. - The package is installed as
scikit-learnbut imported assklearn. - Scikit-Learn provides tools for classification, regression, clustering, preprocessing, and evaluation.
- Most sklearn estimators use the
fit()method to learn from data. - Predictive models commonly use
predict()to generate predictions. - Transformers use
transform()to modify or prepare data. - X usually represents features and y usually represents targets.
- Scikit-Learn works closely with NumPy and Pandas.
- A typical workflow is prepare data → split data → preprocess → train → predict → evaluate.
- Understanding the common sklearn API is more important than memorizing individual algorithms.
Import LinearRegression from sklearn.linear_model, print the installed sklearn version, create a LinearRegression instance, and print the model.
import sklearn from sklearn.linear_model import LinearRegression # Print sklearn version # Create a LinearRegression instance # Print the model # Your code here
Challenge: After creating the model, create a small NumPy dataset, train the model using fit(), and predict the output for a new value.
In Lecture 02, we will install the full data science environment, configure the required Python packages, and verify that everything works correctly before building machine learning projects.
Install and Set Up Scikit-Learn
Install Scikit-Learn and the essential Python data science libraries, create an isolated virtual environment, verify that everything works correctly, understand common installation problems, and learn the import patterns used throughout this course.
- How to install Scikit-Learn with pip
- How to install NumPy, Pandas, and Matplotlib
- How to verify the installation
- How to use virtual environments
- Why virtual environments are important
- Common import patterns for sklearn
- How to check Python and pip versions
- How to troubleshoot common installation errors
- How to test a complete sklearn environment
01. Installing the Full Stack
Scikit-Learn works as part of the wider Python data science ecosystem. Although sklearn provides the machine learning algorithms, real machine learning projects commonly use additional libraries for numerical computation, data manipulation, and visualization.
For this course, we will primarily work with Scikit-Learn, NumPy, Pandas, and Matplotlib.
pip install scikit-learn numpy pandas matplotlib
The command tells pip to download and install the required packages. If the packages are already installed, pip will normally report that the requirements are already satisfied or update them when appropriate.
- NumPy — numerical arrays and mathematical operations
- Pandas — data loading, cleaning, and analysis
- Matplotlib — data visualization and charts
- Scikit-Learn — machine learning algorithms and utilities
02. Checking Python and pip
Before installing packages, it is useful to verify that Python and pip are available from your terminal. This can help identify problems where Python is not installed correctly or where the wrong Python installation is being used.
python --version pip --version
You should see the installed Python version and information about the pip package manager. On some systems, the Python command may be python3 instead.
python3 --version pip3 --version
If your computer has multiple Python installations, always make sure that the pip you use installs packages into the same Python environment that runs your program.
03. Verifying the Installation
After installation, the next step is to verify that Python can successfully import the required libraries.
import sklearn
import numpy as np
import pandas as pd
import matplotlib
print("sklearn:", sklearn.**version**)
print("numpy:", np.**version**)
print("pandas:", pd.**version**)
print("matplotlib:", matplotlib.**version**)
print("All installed!")If this program runs without an import error, the basic environment is working correctly.
Version information is useful when debugging projects. If another developer or a tutorial uses a different library version, knowing your installed versions can help explain differences in behavior.
04. Virtual Environment (Recommended)
A virtual environment creates an isolated Python environment for a project. Packages installed inside the environment are separated from packages installed in other projects.
This is particularly useful for machine learning because different projects may require different versions of libraries.
python -m venv venv
The command creates a directory named venv containing an isolated Python environment.
On Windows, activate it using:
venv\Scripts\activate
On macOS or Linux, use:
source venv/bin/activate
Once the environment is activated, install the project dependencies inside it.
pip install scikit-learn numpy pandas matplotlib
- Keeps project dependencies isolated
- Prevents different projects from interfering with each other
- Makes dependency management easier
- Helps reproduce projects on other computers
- Reduces the risk of breaking the global Python installation
05. Activating and Deactivating the Environment
When you work on a project, activate its virtual environment before running or installing project dependencies.
On Windows:
venv\Scripts\activate
On macOS and Linux:
source venv/bin/activate
When you are finished working on the project, you can leave the virtual environment with:
deactivate
Creating a virtual environment does not automatically activate it. Make sure the correct environment is active before installing packages or running your project.
06. Standard Import Pattern
Once the environment is ready, machine learning programs typically import the specific classes and functions required for the task.
import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error, accuracy_score
Notice that NumPy and Pandas use common aliases, np and pd. Scikit-Learn classes are usually imported directly from their specific modules.
Prefer importing the specific sklearn classes you need instead of importing every sklearn component. This makes your code easier to read and understand.
07. Understanding sklearn Modules
Scikit-Learn is organized into different modules. Each module contains tools designed for a particular part of the machine learning workflow.
sklearn.model_selection— splitting data and selecting modelssklearn.preprocessing— scaling and transforming featuressklearn.linear_model— linear and logistic modelssklearn.tree— decision tree algorithmssklearn.ensemble— ensemble learning algorithmssklearn.neighbors— nearest-neighbor algorithmssklearn.svm— support vector machine algorithmssklearn.cluster— clustering algorithmssklearn.metrics— evaluation metrics
Understanding this organization makes it easier to search the documentation and find the appropriate tool for a particular machine learning task.
08. Testing a Simple sklearn Model
A successful import is useful, but we can perform a stronger test by actually creating and training a model.
import numpy as np
from sklearn.linear_model import LinearRegression
X = np.array([
[1],
[2],
[3],
[4]
])
y = np.array([
2,
4,
6,
8
])
model = LinearRegression()
model.fit(X, y)
prediction = model.predict([[5]])
print("Prediction:", prediction[0])If the program runs successfully and produces a prediction, your environment is capable of executing a basic Scikit-Learn machine learning workflow.
09. Installing Packages with python -m pip
On systems with multiple Python installations, using python -m pip can make it clearer which Python interpreter is being used to run pip.
python -m pip install scikit-learn numpy pandas matplotlib
This tells Python to execute pip as a module. It is often useful when the command pip points to a different Python installation than the one you are using to run your program.
When working with virtual environments, using python -m pip can help ensure that packages are installed for the Python interpreter you intend to use.
10. Checking Installed Packages
pip can display the packages installed in the current environment.
pip list
You can also inspect information about a particular package.
pip show scikit-learn
The command can display information such as the installed version, package location, and dependencies.
pip list— displays installed packagespip show scikit-learn— displays information about Scikit-Learnpip --version— displays pip informationpython --version— displays the Python version
11. Common Installation Errors
Installation problems are common when setting up Python environments. Most errors can be diagnosed by checking the Python version, active environment, pip configuration, and package names.
- ModuleNotFoundError — the required package may not be installed in the active environment.
- Wrong Python environment — pip may have installed the package into a different Python installation.
- Permission errors — the system may not allow modification of a global Python installation.
- Network errors — pip may be unable to download packages.
- Old pip — updating pip may resolve some installation problems.
If you receive an import error after installation, first check whether the environment where the package was installed is the same environment used to run the Python program.
12. Updating pip
Keeping pip reasonably up to date can help with package installation and dependency management.
python -m pip install --upgrade pip
After updating pip, you can install or update your project dependencies normally.
Do not randomly upgrade every package in a production project without considering compatibility. In a real project, dependency versions should be managed deliberately.
13. Creating a Requirements File
When a project uses several Python packages, it is useful to record its dependencies in a requirements.txt file. This makes it easier to recreate the environment on another computer.
numpy pandas matplotlib scikit-learn
Another environment can install these dependencies using:
pip install -r requirements.txt
Requirements files are particularly useful when sharing projects with classmates, teammates, servers, or other development machines.
14. Complete Environment Verification
We can combine the checks from this lecture into one script that verifies the important libraries and performs a simple machine learning operation.
import sklearn
import numpy as np
import pandas as pd
import matplotlib
from sklearn.linear_model import LinearRegression
print("Scikit-Learn:", sklearn.**version**)
print("NumPy:", np.**version**)
print("Pandas:", pd.**version**)
print("Matplotlib:", matplotlib.**version**)
X = np.array([[1], [2], [3]])
y = np.array([2, 4, 6])
model = LinearRegression()
model.fit(X, y)
print("Test prediction:", model.predict([[4]])[0])
print("Environment is working!")This test checks both package imports and a basic sklearn operation. If it runs without errors, your environment is ready for the upcoming machine learning lessons.
15. Best Practices for the Course
- Create a separate virtual environment for your machine learning projects.
- Install packages inside the active environment.
- Use
python -m pipwhen you need to make the Python interpreter explicit. - Keep a
requirements.txtfile for shareable projects. - Check package versions when troubleshooting.
- Use meaningful project folders and keep environments isolated.
- Do not install unnecessary packages into every project.
- Test the environment before beginning a large machine learning project.
A clean environment saves time later. Many machine learning errors are caused not by the algorithm itself but by incorrect packages, incompatible versions, or running code in the wrong Python environment.
16. Key Points
- Install Scikit-Learn with
pip install scikit-learn. - Install NumPy, Pandas, and Matplotlib for the complete basic data science stack.
- Use
python --versionandpip --versionto inspect your environment. - Virtual environments isolate project dependencies.
- Activate the correct virtual environment before installing packages.
- Use
python -m pipwhen you need to ensure pip belongs to a specific Python interpreter. - Verify installations by importing libraries and printing their versions.
- Scikit-Learn is organized into modules such as
linear_model,preprocessing,model_selection, andmetrics. - A simple model test can confirm that the sklearn installation actually works.
- Requirements files make it easier to reproduce project environments.
Create a virtual environment named ml_env, activate it, install Scikit-Learn, NumPy, Pandas, and Matplotlib, and create a Python script that imports all four libraries.
import sklearn
import numpy as np
import pandas as pd
import matplotlib
print("sklearn:", sklearn.**version**)
print("numpy:", np.**version**)
print("pandas:", pd.**version**)
print("matplotlib:", matplotlib.**version**)
# Your code hereChallenge: Import LinearRegression, create a small dataset, train a model using fit(), and make one prediction. Then create a requirements.txt file containing the libraries used by your project.
In Lecture 03, we will learn about NumPy arrays and the numerical data structures used by Scikit-Learn, including shapes, dimensions, indexing, and preparing data for machine learning.
Understand the Machine Learning Workflow
A machine learning model is not created by simply calling one function. Scikit-Learn provides tools that allow you to follow a complete workflow: collect or load data, prepare it, split it into training and testing sets, choose a model, train it, make predictions, evaluate its performance, and improve it when necessary.
01. Understand the Machine Learning Workflow
A typical machine learning project follows a series of steps. Each step has a specific purpose and contributes to building a reliable model.
The general workflow can be represented as:
1. Collect or load data 2. Explore the data 3. Prepare the data 4. Split the data 5. Choose a model 6. Train the model 7. Make predictions 8. Evaluate the model 9. Improve the model
Scikit-Learn provides functions and classes for most of these steps. For example, train_test_split() can divide the data, while classes such as LinearRegression, KNeighborsClassifier, and DecisionTreeClassifier can be used to create machine learning models.
The important idea is that training a model is only one part of the machine learning workflow. A model must also be tested on data that it did not see during training.
Data β Preparation β Split β Model β Train β Predict β Evaluate β Improve
02. Load the Required Libraries
Before building a machine learning model, import the libraries and tools that you need. Scikit-Learn contains many machine learning algorithms and utilities.
import numpy as np import pandas as pd from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error, r2_score
Here, NumPy is useful for numerical operations, while Pandas is commonly used to load and work with tabular data.
From Scikit-Learn, we import train_test_split() for dividing the dataset, LinearRegression for creating a regression model, and evaluation functions for measuring how well the model performs.
You do not need to import every Scikit-Learn module in every project. Import only the tools required for the particular machine learning task.
03. Prepare Features and Targets
Machine learning models usually work with two main types of data: features and a target.
Features are the input values used by the model to make a prediction. The target is the value that the model is trying to predict.
For example, suppose we want to predict a student's exam score from the number of hours they studied.
hours = [[1], [2], [3], [4], [5], [6]] scores = [35, 42, 50, 58, 67, 75]
Here, hours is the feature because it is used as the input to the model. The scores are the target values because they are what we want the model to predict.
Scikit-Learn normally represents features as a two-dimensional structure, even when there is only one feature. That is why the values are written as [[1], [2], [3]] rather than simply [1, 2, 3].
X usually represents the features or input data, while y usually represents the target or output that the model learns to predict.
04. Split the Data
A model should not be evaluated only on the same data that it used for learning. Instead, divide the dataset into a training set and a testing set.
The training data is used to teach the model. The testing data is kept separate and is used later to check how well the model performs on unseen examples.
Scikit-Learn provides the train_test_split() function for this purpose.
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split( hours, scores, test_size=0.2, random_state=42 )
The function returns four values. X_train and y_train contain the training data, while X_test and y_test contain the testing data.
The test_size=0.2 argument means that approximately 20% of the available data is reserved for testing, while the remaining 80% is used for training.
The random_state=42 argument makes the split reproducible. This means that running the same code again will produce the same split.
Do not train your model on the test data and then use that same data to claim that the model performs well. The test set should represent data that the model has not seen during training.
05. Choose a Machine Learning Model
After preparing and splitting the data, choose an algorithm that matches the problem you are solving.
Different machine learning problems require different types of models.
| Problem | Example | Possible Scikit-Learn Model |
|---|---|---|
| Regression | Predict house price | LinearRegression |
| Classification | Predict spam or not spam | LogisticRegression |
| Classification | Classify flowers | KNeighborsClassifier |
| Classification | Classify customers | DecisionTreeClassifier |
Choosing the correct model depends on the type of problem, the dataset, the number of features, the amount of data, and the expected output.
For this example, we are predicting a numerical exam score, so regression is appropriate. We can use LinearRegression.
06. Create and Train the Model
Once a model has been selected, create an instance of the model and use the fit() method to train it.
from sklearn.linear_model import LinearRegression model = LinearRegression() model.fit(X_train, y_train)
The first line creates a LinearRegression object. The second line calls fit(), which allows the model to learn the relationship between the input features and the target values.
The model examines the training data and learns parameters that can later be used to make predictions.
In Scikit-Learn, the fit() method is one of the most important methods you will encounter. Many Scikit-Learn models follow the same general pattern:
model = SomeModel() model.fit(X_train, y_train)
fit() means learn from the training data. It does not mean that the model is permanently correct. The model still needs to be evaluated using data that was not used for training.
07. Make Predictions
After training, use the predict() method to generate predictions.
predictions = model.predict(X_test) print(predictions)
The model receives the test features stored in X_test and produces predicted values.
For example, if X_test contains the number of hours studied by students that the model has never seen before, the model will use what it learned during training to estimate their exam scores.
The predicted values are stored in the predictions variable.
Prediction can also be performed on completely new data.
new_student = [[7]] predicted_score = model.predict(new_student) print(predicted_score)
Here, the model receives a new input representing seven hours of study and estimates the corresponding exam score.
08. Evaluate the Model
A prediction alone does not tell us whether a model is good. We need evaluation metrics to measure its performance.
For regression problems, common metrics include Mean Squared Error (MSE) and RΒ² score.
from sklearn.metrics import mean_squared_error, r2_score mse = mean_squared_error(y_test, predictions) r2 = r2_score(y_test, predictions) print(βMSE:β, mse) print(βRΒ²:β, r2)
mean_squared_error() measures the average squared difference between the actual values and predicted values. A lower MSE generally indicates that the predictions are closer to the actual values.
r2_score() measures how well the model explains the variation in the target values. A value closer to 1 generally indicates a stronger fit for a regression model.
The correct evaluation metric depends on the type of machine learning problem. Classification models use metrics such as accuracy, precision, recall, and F1-score, while regression models commonly use metrics such as MSE, MAE, and RΒ².
09. Put the Complete Workflow Together
Now we can combine the major steps into one complete Scikit-Learn program.
from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error, r2_score hours = [[1], [2], [3], [4], [5], [6]] scores = [35, 42, 50, 58, 67, 75] X_train, X_test, y_train, y_test = train_test_split( hours, scores, test_size=0.2, random_state=42 ) model = LinearRegression() model.fit(X_train, y_train) predictions = model.predict(X_test) mse = mean_squared_error(y_test, predictions) r2 = r2_score(y_test, predictions) print(βPredictions:β, predictions) print(βMSE:β, mse) print(βRΒ²:β, r2)
This program demonstrates the complete basic workflow. First, the data is created. Then it is divided into training and testing sets. A regression model is created and trained. The trained model makes predictions on the test data, and the predictions are finally evaluated.
Learning this workflow is more important than memorizing individual algorithms because many Scikit-Learn models use a similar interface.
10. Improve the Model
If the model does not perform well, the workflow does not simply end. Machine learning is usually an iterative process.
You may need to improve the data, select different features, preprocess the values, choose another algorithm, tune hyperparameters, or collect more training data.
A simplified improvement cycle looks like this:
Prepare data
β
Train model
β
Evaluate model
β
Analyze results
β
Improve data or model
β
Train again
β
Evaluate again
For example, a model may perform poorly because the features are not useful, the data contains missing values, the model is too simple, or the model is overfitting the training data.
Scikit-Learn provides additional tools for preprocessing, feature selection, cross-validation, and hyperparameter tuning that can be introduced as you progress through machine learning.
11. The Scikit-Learn Estimator Pattern
One of the most useful things to understand about Scikit-Learn is that many algorithms follow a consistent interface.
A typical model follows this pattern:
model = Model() model.fit(X_train, y_train) predictions = model.predict(X_test)
This pattern makes Scikit-Learn easier to learn. Once you understand how fit() and predict() work, you can apply the same basic idea to many different algorithms.
Some models also provide methods such as score(), transform(), or fit_transform(). These methods become especially important when working with preprocessing tools and machine learning pipelines.
- Load or collect the dataset.
- Identify the features and target.
- Explore and understand the data.
- Clean and prepare the data.
- Split the data into training and testing sets.
- Choose an appropriate machine learning algorithm.
- Create the model.
- Train it using
fit(). - Generate predictions using
predict(). - Evaluate the predictions with suitable metrics.
- Improve the model when necessary.
- Repeat the process until the model meets the required goal.
Never allow information from the test set to influence the training process. This can produce unrealistically good evaluation results and make the model appear better than it really is.
As you learn preprocessing and scaling, you will see why operations such as fitting a StandardScaler should normally be learned from the training data rather than the test data.
Create a small machine learning program that predicts a numerical value. You can use a dataset such as study hours and exam scores, house size and house price, or advertising budget and sales.
Your program should:
- Create or load a dataset.
- Separate the features from the target.
- Split the data into training and testing sets.
- Create an appropriate Scikit-Learn model.
- Train the model using
fit(). - Generate predictions using
predict(). - Evaluate the model using at least one suitable metric.
- Make a prediction for a new input value.
When you finish, explain each stage of your program in your own words. You should be able to answer: What is the input? What is the target? Which data was used for training? Which data was used for testing? What did the model learn? How well did it perform?
Work with Datasets in Scikit-Learn
Machine learning models need data to learn from. Before training a model, you need to know how to obtain a dataset, load it into Python, understand its structure, identify features and targets, and prepare it for the machine learning workflow. Scikit-Learn provides several built-in datasets for learning and experimentation, while Pandas makes it easy to load real-world datasets from files such as CSV.
01. What Is a Dataset?
A dataset is a collection of data that can be used to analyze a problem or train a machine learning model. A dataset usually contains multiple observations, where each observation represents one example.
For example, a student dataset might contain information about study hours, attendance, assignments, and exam scores.
Hours Attendance Assignments Score 2 75 6 48 4 82 8 61 6 90 9 74 8 95 10 88
Each row represents one student, while each column represents a variable or feature.
If our goal is to predict the exam score, Hours, Attendance, and Assignments could be features, while Score would be the target.
Rows usually represent individual observations or examples, while columns represent features, variables, or the target.
02. Features and Target
Before training a machine learning model, separate the dataset into input features and the target value.
Features are the information that the model uses to make a prediction. The target is the value that the model is trying to predict.
import pandas as pd
data = pd.DataFrame({
βHoursβ: [2, 4, 6, 8],
βAttendanceβ: [75, 82, 90, 95],
βScoreβ: [48, 61, 74, 88]
})
X = data[[βHoursβ, βAttendanceβ]]
y = data[βScoreβ]
print(X)
print(y)
Here, X contains the input features, while y contains the target values.
Notice that X uses double square brackets because it contains multiple columns and therefore remains a two-dimensional DataFrame. The target y is a single Pandas Series.
This separation is an important step because Scikit-Learn expects the model to receive the features separately from the values it should learn to predict.
03. Use Built-in Scikit-Learn Datasets
Scikit-Learn provides several built-in datasets that are useful for learning machine learning concepts without downloading external files.
Some commonly used datasets include:
| Dataset | Purpose | Typical Task |
|---|---|---|
load_iris() |
Flower measurements | Classification |
load_digits() |
Handwritten digits | Classification |
load_breast_cancer() |
Medical diagnostic data | Classification |
load_diabetes() |
Diabetes-related measurements | Regression |
load_wine() |
Wine measurements | Classification |
These datasets are particularly useful when learning because they allow you to practice machine learning without first finding and downloading a dataset yourself.
04. Load the Iris Dataset
The Iris dataset is one of the most commonly used beginner datasets in machine learning. It contains measurements of iris flowers and information about their species.
Scikit-Learn provides it through the load_iris() function.
from sklearn.datasets import load_iris iris = load_iris() print(iris)
The returned object contains several pieces of information about the dataset, including the input data, target values, feature names, and target names.
Instead of printing the entire dataset, you can access individual parts of it.
from sklearn.datasets import load_iris iris = load_iris() print(iris.data) print(iris.target) print(iris.feature_names) print(iris.target_names)
iris.data contains the input features, while iris.target contains the numerical class labels.
iris.feature_names provides the names of the features, and iris.target_names tells us which classes the numerical target values represent.
05. Understand the Dataset Shape
Before training a model, it is useful to know how much data you have and how many features are available.
NumPy arrays and Pandas DataFrames provide the shape attribute for this purpose.
from sklearn.datasets import load_iris iris = load_iris() X = iris.data y = iris.target print(βFeatures shape:β, X.shape) print(βTarget shape:β, y.shape)
The shape is usually displayed as a pair such as (150, 4). The first number represents the number of observations, while the second represents the number of features.
For the Iris dataset, there are 150 examples and 4 input features.
Checking the shape helps you quickly understand your dataset and can reveal problems such as missing columns, unexpected dimensions, or incorrectly formatted input data.
06. Inspect the Data
Do not immediately train a model after loading a dataset. First inspect the data so that you understand what you are working with.
When using Pandas, common inspection methods include head(), info(), describe(), and shape.
import pandas as pd data = pd.read_csv(βstudents.csvβ) print(data.head()) print(data.shape) print(data.info()) print(data.describe())
head() displays the first few rows, which gives you a quick look at the data.
shape tells you how many rows and columns are present.
info() provides information about the columns, data types, and non-null values.
describe() generates statistical information for numerical columns, such as the mean, minimum, maximum, and quartiles.
These checks are useful because real-world datasets are rarely perfectly prepared for machine learning.
07. Load a CSV Dataset
Real machine learning projects often use datasets stored in files. One of the most common formats is CSV, which stands for Comma-Separated Values.
Pandas provides the read_csv() function for loading CSV files.
import pandas as pd data = pd.read_csv(βstudents.csvβ) print(data.head())
The string "students.csv" represents the location of the CSV file. If the file is in the same folder as your Python program, you can use just the filename.
If the file is stored somewhere else, you can provide a file path.
data = pd.read_csv("data/students.csv")
After loading the file, data becomes a Pandas DataFrame that can be inspected and prepared for machine learning.
08. Select Features and Target from a CSV
Once the CSV file has been loaded, select the columns that will be used as features and identify the target column.
Suppose the CSV contains the following columns:
Hours,Attendance,Assignments,Score 2,75,6,48 4,82,8,61 6,90,9,74 8,95,10,88
If we want to predict Score, the other columns can be used as input features.
import pandas as pd data = pd.read_csv(βstudents.csvβ) X = data[[βHoursβ, βAttendanceβ, βAssignmentsβ]] y = data[βScoreβ] print(X) print(y)
The variable X now contains the information used to make predictions, while y contains the values the model should learn to predict.
The target should be the value you actually want the model to predict. Do not accidentally include the target column inside X, because this can cause data leakage and produce misleading results.
09. Convert Data into NumPy Arrays
Scikit-Learn works naturally with NumPy arrays and Pandas DataFrames. In many situations, you can pass a DataFrame directly to a Scikit-Learn model.
For example:
import pandas as pd data = pd.read_csv(βstudents.csvβ) X = data[[βHoursβ, βAttendanceβ]] y = data[βScoreβ] print(type(X)) print(type(y))
You can also convert Pandas data into NumPy arrays when needed.
X_array = X.to_numpy() y_array = y.to_numpy() print(X_array) print(y_array)
However, conversion is not always necessary. Most Scikit-Learn estimators can work directly with Pandas DataFrames and Series.
10. Load Different Built-in Datasets
Once you understand the Iris dataset, you can experiment with other datasets provided by Scikit-Learn.
For example, the digits dataset contains images represented as numerical features.
from sklearn.datasets import load_digits digits = load_digits() X = digits.data y = digits.target print(βFeatures:β, X.shape) print(βTargets:β, y.shape)
You can also load the wine dataset:
from sklearn.datasets import load_wine wine = load_wine() X = wine.data y = wine.target print(βFeatures:β, X.shape) print(βTargets:β, y.shape)
Experimenting with different datasets helps you understand that machine learning models can work with many types of numerical data and different prediction problems.
11. Loading Data Is Only the Beginning
Loading a dataset does not automatically make it ready for a machine learning model. Real datasets may contain missing values, duplicate rows, text values, categorical variables, extreme values, or features with very different scales.
A typical preparation process may look like this:
Load dataset
β
Inspect dataset
β
Handle missing values
β
Select features
β
Select target
β
Encode categorical data
β
Scale features when necessary
β
Split into training and testing data
β
Train the model
Each dataset may require a different preparation process. The goal is to convert raw data into a form that a machine learning algorithm can use effectively.
Never assume that a dataset is ready just because Python successfully loaded it. Always inspect the structure, data types, missing values, and target before training a model.
12. Build a Simple Dataset-to-Model Workflow
We can now combine data loading, feature selection, splitting, training, and prediction into one simple example.
import pandas as pd from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression data = pd.read_csv(βstudents.csvβ) X = data[[βHoursβ, βAttendanceβ]] y = data[βScoreβ] X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) model = LinearRegression() model.fit(X_train, y_train) predictions = model.predict(X_test) print(βPredictions:β, predictions)
This example demonstrates an important connection between data loading and the machine learning workflow. The dataset is first loaded using Pandas, then divided into features and a target, and finally passed into the Scikit-Learn training process.
Understanding this connection will make it much easier to work with larger real-world datasets later.
- Always inspect a dataset after loading it.
- Check the number of rows and columns.
- Check the data types of each column.
- Look for missing values.
- Identify the target column before selecting features.
- Keep features and target separate.
- Use built-in Scikit-Learn datasets when practicing algorithms.
- Use Pandas for convenient loading and inspection of CSV data.
- Do not assume raw data is ready for training.
- Prepare the data before giving it to a machine learning model.
Use one of Scikit-Learn's built-in datasets such as Iris, Digits, Wine, or Diabetes.
Your program should:
- Import and load the dataset.
- Display the feature names if they are available.
- Display the target names if they are available.
- Display the shape of the feature data.
- Display the shape of the target data.
- Print the first few examples.
- Separate the features into
Xand the target intoy.
After completing the exercise, explain what each row and column represents. Then identify whether the dataset is more suitable for a classification or regression problem.
Train / Test Split
Learn why splitting your data is essential, how training and testing data work, and how to use train_test_split correctly to evaluate machine learning models and avoid misleading results.
- Why machine learning data must be divided into training and testing sets
- The difference between training data and testing data
- How to use
train_test_split() - How to choose an appropriate test size
- Why
random_stateis useful - How
stratifypreserves class proportions - How data splitting helps measure generalization
- Common mistakes to avoid when splitting datasets
01. Why Split Data?
If you train and evaluate a model using exactly the same data, the model may appear to perform extremely well because it has already seen those examples. This does not tell us how well the model will perform on completely new data.
This problem is closely related to overfitting. An overfitted model learns the training examples too closely, including patterns that may not generalize to unseen data.
By separating the dataset into training and testing portions, we can train the model on one portion and evaluate it on data that was kept unseen during training.
Training data teaches the model. Testing data checks whether the model learned something that generalizes to new examples.
from sklearn.model_selection import train_test_split
from sklearn.datasets import load_iris
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42,
stratify=y
)
print("Train:", X_train.shape)
print("Test:", X_test.shape)02. Training Data vs Testing Data
The training set is the portion of the dataset used to learn the relationships between features and targets. The testing set is kept separate until evaluation.
Training Set
Used by the model to learn patterns and relationships between features and targets.
Testing Set
Contains examples that the model should not see during training.
Generalization
Describes how well a trained model performs on unseen data.
Evaluation
Measures model performance using the test data and appropriate metrics.
For example, if a dataset contains 1,000 samples and we use an 80/20 split, approximately 800 samples are used for training and 200 samples are reserved for testing.
03. Using train_test_split()
Scikit-Learn provides the train_test_split() function in sklearn.model_selection. It randomly divides one or more arrays into training and testing subsets.
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2 ) print(X_train.shape) print(X_test.shape) print(y_train.shape) print(y_test.shape)
When both X and y are supplied, Scikit-Learn keeps their corresponding rows aligned. This is extremely important because each feature row must remain associated with the correct target.
04. Understanding test_size
The test_size parameter controls how much data is placed into the testing set.
# 20% of the data for testing
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.2
)
# 30% of the data for testing
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.3
)
# Exactly 100 samples for testing
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=100
)test_size=0.2— 20% testing and approximately 80% training.test_size=0.25— 25% testing and approximately 75% training.test_size=0.3— 30% testing and approximately 70% training.test_size=100— exactly 100 samples for testing.
There is no universal test size that is correct for every project. The appropriate choice depends on the dataset size, problem, and evaluation strategy.
05. Understanding random_state
By default, the split is random. Running the same program multiple times can therefore produce different training and testing samples.
The random_state parameter allows you to control the randomness. Using the same value produces the same split each time, which makes experiments easier to reproduce.
from sklearn.model_selection import train_test_split X_train1, X_test1, y_train1, y_test1 = train_test_split( X, y, test_size=0.2, random_state=42 ) X_train2, X_test2, y_train2, y_test2 = train_test_split( X, y, test_size=0.2, random_state=42 ) print((X_train1 == X_train2).all())
The number 42 has no special machine learning meaning. It is simply a commonly used seed value. You can use another integer as long as you document it when reproducibility matters.
06. Understanding stratify
For classification problems, different classes may have different numbers of examples. A random split can sometimes produce slightly different class proportions between training and testing sets.
The stratify parameter helps preserve the class distribution when splitting the data.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)
print("Training classes:", y_train)
print("Testing classes:", y_test)Suppose a classification dataset contains 70% class A and 30% class B. Stratification attempts to maintain approximately the same class proportions in both the training and testing sets.
07. Checking the Split
After splitting the data, it is good practice to check the shapes of all four resulting arrays. This confirms that the split happened as expected.
print("X_train:", X_train.shape)
print("X_test :", X_test.shape)
print("y_train:", y_train.shape)
print("y_test :", y_test.shape)
print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))For the Iris dataset, there are 150 samples. With test_size=0.2, the test set contains 30 samples and the training set contains 120 samples.
08. Splitting Multiple Arrays
One important advantage of train_test_split() is that it can split multiple related arrays at the same time. This is how we keep features and targets correctly matched.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)
print("Training features:", X_train.shape)
print("Training targets:", y_train.shape)
print("Testing features:", X_test.shape)
print("Testing targets:", y_test.shape)Never independently shuffle X and y. Doing so can destroy the relationship between a feature row and its correct target. Use train_test_split(X, y, ...) so Scikit-Learn keeps them synchronized.
09. Train / Test Split with a CSV Dataset
The same technique works when the data comes from your own CSV file. First load the data using Pandas, separate the features and target, and then split them.
import pandas as pd
from sklearn.model_selection import train_test_split
df = pd.read_csv("data.csv")
X = df.drop("target", axis=1)
y = df["target"]
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)
print("Training data:", X_train.shape)
print("Testing data:", X_test.shape)10. Why We Should Not Train on Test Data
The test set should represent data that the model has never seen during training. If test data is accidentally used during training, the evaluation can become overly optimistic.
This is one form of data leakage. Data leakage occurs when information that should be unavailable to the model during training finds its way into the training process.
- Do not train the model using
X_testandy_test. - Do not use test data to choose the best model.
- Do not calculate preprocessing statistics from the entire dataset before splitting.
- Keep the test set isolated until final evaluation.
11. Train / Test Split and Preprocessing
When preprocessing data, you should generally split the data first and then fit preprocessing transformations using the training data. The transformation can then be applied to the test data.
from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) scaler = StandardScaler() X_train = scaler.fit_transform(X_train) X_test = scaler.transform(X_test)
Use fit_transform() on training data because the scaler must learn its parameters from the training set. Use only transform() on test data so information from the test set does not influence the learned transformation.
12. Train / Test Split vs Validation Data
For simple projects, a train/test split may be enough. However, larger machine learning workflows often divide data into three sets: training, validation, and testing.
Training Set
Used to train the model and learn its parameters.
Validation Set
Used during development to compare models and tune hyperparameters.
Testing Set
Used at the end to estimate performance on unseen data.
Cross-Validation
Provides another way to evaluate models using multiple training and validation splits.
We will study validation and cross-validation in more detail later in the course.
13. Complete Train / Test Example
The following example combines the main concepts from this lecture into a complete workflow.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
# Load data
X, y = load_iris(return_X_y=True)
# Split into training and testing sets
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)
# Display sizes
print("Training features:", X_train.shape)
print("Testing features:", X_test.shape)
print("Training targets:", y_train.shape)
print("Testing targets:", y_test.shape)14. Common Mistakes
- Training and testing on the exact same data.
- Forgetting to separate
Xandy. - Using the test set during model training.
- Using test data to repeatedly tune the model.
- Ignoring class imbalance in classification problems.
- Forgetting
random_statewhen reproducible experiments are required. - Fitting preprocessing transformations on the complete dataset before splitting.
- Using an extremely small test set that does not provide a useful evaluation.
15. Key Points
- Training data is used to teach the model.
- Testing data is used to evaluate the model on unseen examples.
train_test_split()is provided bysklearn.model_selection.test_sizecontrols the size of the test set.random_statemakes random splitting reproducible.stratify=yhelps preserve class proportions in classification tasks.- Never allow test data to influence the training process.
- Data preprocessing should be fitted using training data only.
- A proper split helps provide a more realistic estimate of model generalization.
Load the Iris dataset and split it into 80% training data and 20% testing data using random_state=0 and stratify=y. Print the shapes of all four arrays.
from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split X, y = load_iris(return_X_y=True) # Split the data # Your code here # Print all four shapes # X_train # X_test # y_train # y_test
In Lecture 06, we will learn about Data Preprocessing, including handling missing values, encoding categorical data, scaling numerical features, and preparing raw data for machine learning models.
Feature Scaling and Normalization
Learn how to scale your data so all features contribute appropriately to model training, understand why different numerical ranges can affect machine learning algorithms, and apply common Scikit-Learn scaling techniques correctly.
- Why feature scaling is important in machine learning
- How StandardScaler works
- How MinMaxScaler works
- The difference between standardization and normalization
- Why scaling must be fitted only on training data
- How scaling prevents data leakage
- Which machine learning algorithms are affected by feature scale
- How to choose an appropriate scaling technique
01. Why Feature Scaling?
Machine learning datasets often contain features measured on completely different scales. For example, one feature might represent age from 0 to 100, while another feature represents salary from 10,000 to 1,000,000.
If an algorithm depends on distances, magnitudes, or numerical optimization, a feature with much larger values can have a disproportionate influence on the model.
Imagine two features: age ranges from 18 to 80, while income ranges from 20,000 to 200,000. Without scaling, the numerical magnitude of income is much larger than age, even though both may be equally important.
Feature scaling transforms numerical columns to a more comparable scale while preserving the useful information contained in the data.
02. Standardization vs Normalization
Two common approaches to feature scaling are standardization and normalization. Although these terms are sometimes used interchangeably, they describe different transformations.
Standardization
Transforms features so they generally have a mean near 0 and a standard deviation near 1.
Normalization
Often refers to scaling values to a fixed range such as 0 to 1, depending on the technique being used.
StandardScaler
Uses the mean and standard deviation of the training data.
MinMaxScaler
Maps feature values to a specified range, commonly 0 to 1.
03. StandardScaler
StandardScaler is one of the most commonly used preprocessing tools in Scikit-Learn. It standardizes each feature using statistics calculated from the training data.
The transformation is based on the formula:
z = (x - mean) / standard deviation
After standardization, a feature will typically have a mean close to 0 and a standard deviation close to 1 on the data used to fit the scaler.
from sklearn.preprocessing import StandardScaler
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42
)
scaler = StandardScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
print("Mean:", X_train_s.mean(axis=0).round(2))
print("Std: ", X_train_s.std(axis=0).round(2))Notice that the training data uses fit_transform(), while the test data uses only transform(). This distinction is extremely important.
04. Understanding fit(), transform(), and fit_transform()
Scikit-Learn preprocessing objects follow the same API pattern used throughout the library.
scaler = StandardScaler() # Learn statistics from training data scaler.fit(X_train) # Transform the training data X_train_s = scaler.transform(X_train) # Transform new data using the same statistics X_test_s = scaler.transform(X_test)
fit()— learns parameters such as the mean and standard deviation.transform()— applies the learned transformation.fit_transform()— performs both operations in one call.
The common training pattern is therefore:
scaler.fit(X_train) X_train_scaled = scaler.transform(X_train) X_test_scaled = scaler.transform(X_test)
05. MinMaxScaler
MinMaxScaler transforms features into a specified range. The default range is from 0 to 1.
This can be useful when you want your features to have a predictable minimum and maximum value.
from sklearn.preprocessing import MinMaxScaler
scaler = MinMaxScaler()
X_train_s = scaler.fit_transform(X_train)
X_test_s = scaler.transform(X_test)
print("Min:", X_train_s.min(axis=0))
print("Max:", X_train_s.max(axis=0))The transformation can be represented conceptually as:
x_scaled = (x - x_min) / (x_max - x_min)
With the default configuration, values are mapped approximately into the range 0 to 1 based on the training data.
06. Choosing a Feature Range
MinMaxScaler allows you to choose a custom output range instead of using 0 to 1.
from sklearn.preprocessing import MinMaxScaler scaler = MinMaxScaler(feature_range=(-1, 1)) X_train_s = scaler.fit_transform(X_train) X_test_s = scaler.transform(X_test) print(X_train_s.min(axis=0)) print(X_train_s.max(axis=0))
This can be useful for algorithms or applications where a particular numerical range is desirable.
07. Why Scaling Helps Some Algorithms
Not every machine learning algorithm is equally affected by feature scale. Algorithms that use distances, gradients, or regularization are often particularly sensitive to differently scaled features.
- K-Nearest Neighbors (KNN)
- Support Vector Machines (SVM)
- Logistic Regression
- Linear Regression with regularization
- Ridge and Lasso regression
- Neural networks
- K-Means clustering
- Principal Component Analysis (PCA)
For example, KNN calculates distances between observations. If one feature has much larger numerical values than another, it can dominate the distance calculation.
08. Algorithms That Often Do Not Require Scaling
Some algorithms are much less sensitive to the numerical scale of individual features. Tree-based algorithms typically make decisions using thresholds rather than distance calculations.
- Decision Trees
- Random Forests
- Gradient-boosted tree models
- Many other tree-based methods
This does not mean scaling is harmful in every situation. It means that scaling is generally not necessary for the core decision-making mechanism of these models.
09. Critical Rule — No Data Leakage
Always fit_transform() on training data only, then transform() on the test set. Fitting on the test set causes data leakage and can give falsely optimistic evaluation results.
scaler = StandardScaler() # Correct X_train_scaled = scaler.fit_transform(X_train) # Correct X_test_scaled = scaler.transform(X_test)
The scaler learns its mean and standard deviation from X_train. The same learned values are then used to transform X_test.
scaler = StandardScaler() X_train_scaled = scaler.fit_transform(X_train) # Do NOT do this X_test_scaled = scaler.fit_transform(X_test)
The second fit_transform() causes the scaler to learn new statistics from the test data. The test set should remain unseen during the fitting process.
10. Scaling After Train/Test Split
The correct order of operations is important. You should normally split the data before fitting the scaler.
from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler # 1. Split the data X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # 2. Create scaler scaler = StandardScaler() # 3. Fit only on training data X_train = scaler.fit_transform(X_train) # 4. Transform test data X_test = scaler.transform(X_test)
Load Data → Split Data → Fit Scaler on Training Data → Transform Training Data → Transform Test Data → Train Model
11. Comparing StandardScaler and MinMaxScaler
Both scalers change the numerical scale of features, but they do so differently.
StandardScaler
Centers data around zero and scales according to standard deviation.
MinMaxScaler
Maps features to a chosen range, commonly 0 to 1.
Standardization
Useful when features have different units and a zero-centered distribution is desirable.
Min-Max Scaling
Useful when a bounded numerical range is important for the application.
12. Inspecting the Scaler
After fitting a scaler, Scikit-Learn stores the learned parameters inside the scaler object. These values can be inspected to understand what the preprocessing step learned from the training data.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
print("Means:", scaler.mean_)
print("Scales:", scaler.scale_)These values are learned from the training data and are then reused whenever new data is transformed using the same scaler.
13. Scaling a New Sample
Once a scaler has been fitted, the same scaler can be used to transform new observations before sending them to a trained model.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train_scaled = scaler.fit_transform(X_train)
new_sample = X_test[:1]
new_sample_scaled = scaler.transform(new_sample)
print("Original:", new_sample)
print("Scaled:", new_sample_scaled)Do not create and fit a brand-new scaler for every new sample. Reuse the scaler that was fitted during model training so the new data is transformed using the same rules.
14. Complete Scaling Example
The following example combines dataset loading, train/test splitting, scaling, and a simple model into one workflow.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.2,
random_state=42,
stratify=y
)
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
model = KNeighborsClassifier(n_neighbors=3)
model.fit(X_train, y_train)
predictions = model.predict(X_test)
print("Accuracy:", accuracy_score(y_test, predictions))15. Common Scaling Mistakes
- Fitting the scaler on the entire dataset before splitting.
- Calling
fit_transform()on the test data. - Creating a different scaler for training and testing data.
- Forgetting to scale new prediction data using the same fitted scaler.
- Assuming every algorithm requires feature scaling.
- Ignoring data leakage during preprocessing.
- Scaling categorical values without first choosing an appropriate encoding strategy.
16. Key Points
- Feature scaling puts numerical features on a more comparable scale.
StandardScalerstandardizes features using the mean and standard deviation.MinMaxScalermaps features into a specified range.- Distance-based and gradient-based algorithms often benefit from scaling.
- Many tree-based algorithms are less sensitive to feature scale.
- Always fit preprocessing objects using training data only.
- Use
fit_transform()on training data. - Use
transform()on test and future data. - Incorrect scaling can cause data leakage.
- The same fitted scaler should be reused when transforming new observations.
Load Iris, split the data into 80% training and 20% testing sets, apply StandardScaler correctly, and print the mean and standard deviation of the scaled training set.
from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler X, y = load_iris(return_X_y=True) # Split the data # Your code here # Create the scaler # Your code here # Scale training and testing data # Your code here # Print the mean and standard deviation # Your code here
In Lecture 07, we will learn about Encoding Categorical Data, including how to convert text-based categories into numerical representations that machine learning models can understand.
Project 1: Iris Flower Classifier
Apply everything from Unit I to build your first complete classification model with Scikit-Learn. You will load real data, explore it, split it correctly, scale the features, train a KNN model, make predictions, and evaluate its performance.
- Load and explore the Iris dataset
- Understand features and target labels
- Split the dataset into training and testing sets
- Scale numerical features with StandardScaler
- Understand how K-Nearest Neighbors works
- Train a KNN classification model
- Make predictions on unseen test data
- Evaluate the model using accuracy
- Generate a classification report
- Test the trained model on new flower measurements
01. Project Overview
In this project, we will build a machine learning model that can identify the species of an Iris flower from measurements of its physical characteristics.
The Iris dataset contains measurements of iris flowers belonging to three species: setosa, versicolor, and virginica.
Sepal Length
Length of the flower's sepal measured in centimeters.
Sepal Width
Width of the flower's sepal measured in centimeters.
Petal Length
Length of the flower's petal measured in centimeters.
Petal Width
Width of the flower's petal measured in centimeters.
These four measurements become our features, while the flower species becomes our target.
02. The Machine Learning Pipeline
This project combines the main concepts covered in Unit I into one complete workflow.
Load Data → Explore Data → Split Data → Scale Features → Create Model → Train Model → Predict → Evaluate
Each step has a specific purpose. Keeping these steps separate makes machine learning projects easier to understand, debug, and maintain.
03. Loading the Iris Dataset
Scikit-Learn provides the Iris dataset through load_iris(). We can load it directly without downloading a CSV file.
from sklearn.datasets import load_iris
iris = load_iris()
print("Feature names:", iris.feature_names)
print("Target names:", iris.target_names)
print("Data shape:", iris.data.shape)
print("Target shape:", iris.target.shape)The dataset contains 150 samples and 4 numerical features. The target contains three possible flower classes.
04. Separating X and y
Machine learning models generally use X to represent input features and y to represent target values.
iris = load_iris()
X = iris.data
y = iris.target
print("X shape:", X.shape)
print("y shape:", y.shape)
print("First X row:", X[0])
print("First y value:", y[0])- X contains the four flower measurements.
- y contains the numerical class labels.
X.shapeis(150, 4).y.shapeis(150,).
05. Splitting the Dataset
We should not train and evaluate our model using exactly the same examples. Instead, we reserve part of the dataset for testing.
from sklearn.model_selection import train_test_split
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)
print("Training samples:", len(X_train))
print("Testing samples:", len(X_test))With 150 total samples and a 20% test size, approximately 120 samples are used for training and 30 samples are reserved for testing.
stratify=y helps maintain the class proportions when creating the training and testing sets. This is especially useful for classification problems.
06. Scaling the Features
KNN is a distance-based algorithm. This means the numerical scale of the features can affect how distances between flowers are calculated.
We therefore use StandardScaler to standardize the features.
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
print("Scaled training data:")
print(X_train[:5])Use fit_transform() on the training data and only transform() on the test data. Never fit the scaler separately on the test set.
07. Understanding K-Nearest Neighbors
K-Nearest Neighbors (KNN) is a supervised machine learning algorithm that classifies a new observation based on nearby training observations.
When a new flower is given to the model, KNN looks at the closest training examples and uses their classes to determine the predicted class.
If most of the five nearest flowers belong to the versicolor class, the model will generally predict the new flower as versicolor.
08. Choosing n_neighbors
The n_neighbors parameter determines how many nearby training samples KNN considers when making a prediction.
from sklearn.neighbors import KNeighborsClassifier model_3 = KNeighborsClassifier(n_neighbors=3) model_5 = KNeighborsClassifier(n_neighbors=5) model_10 = KNeighborsClassifier(n_neighbors=10)
Using a very small value can make the model highly sensitive to individual training examples, while a larger value considers a broader neighborhood.
We will use n_neighbors=5. Later in the course, you will learn how to systematically choose model hyperparameters instead of selecting them manually.
09. Creating and Training the Model
Once our data has been prepared, we create a KNN classifier and train it using the training data.
from sklearn.neighbors import KNeighborsClassifier
model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)
print("Model trained successfully!")The fit() method gives the model access to the training features and their corresponding target labels.
10. Making Predictions
After training, we can use the model to predict the classes of flowers in the test set.
y_pred = model.predict(X_test)
print("Predicted:", y_pred)
print("Actual: ", y_test)The predicted values are compared with the actual target values to determine how well the model performed.
11. Measuring Accuracy
Accuracy measures the proportion of predictions that were correct.
from sklearn.metrics import accuracy_score
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
print("Accuracy percentage:", accuracy * 100, "%")If the model correctly predicts 29 out of 30 test samples, its accuracy is 29 / 30 = approximately 96.7%.
12. Classification Report
Accuracy gives us one overall measurement, but classification problems often require more detailed evaluation. Scikit-Learn provides classification_report() for this purpose.
from sklearn.metrics import classification_report print( classification_report( y_test, y_pred, target_names=iris.target_names ) )
The report includes metrics such as precision, recall, and F1-score for each class.
- Precision — how many predicted positives were actually correct.
- Recall — how many actual positives were successfully identified.
- F1-score — combines precision and recall into a single metric.
- Support — number of actual samples belonging to each class.
13. Complete Project Code
Now we can combine everything into one complete Python program.
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score, classification_report
# Load dataset
iris = load_iris()
X, y = iris.data, iris.target
print("Shape:", X.shape)
print("Classes:", iris.target_names)
# Split data
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42,
stratify=y
)
# Scale features
scaler = StandardScaler()
X_train = scaler.fit_transform(X_train)
X_test = scaler.transform(X_test)
# Create model
model = KNeighborsClassifier(n_neighbors=5)
# Train model
model.fit(X_train, y_train)
# Predict
y_pred = model.predict(X_test)
# Evaluate
accuracy = accuracy_score(y_test, y_pred)
print("Accuracy:", accuracy)
print(classification_report(
y_test,
y_pred,
target_names=iris.target_names
))14. Making a Prediction for a New Flower
One of the most useful parts of machine learning is making predictions on new data. Suppose we receive measurements for a flower that was not part of the original training dataset.
new_flower = [[5.1, 3.5, 1.4, 0.2]]
new_flower_scaled = scaler.transform(new_flower)
prediction = model.predict(new_flower_scaled)
print("Predicted class:", prediction[0])
print("Predicted species:", iris.target_names[prediction[0]])The new flower must be transformed using the same scaler that was fitted on the training data. Do not create and fit a new scaler for the new sample.
15. Understanding the Complete Pipeline
At this point, the entire project can be understood as a sequence of connected steps.
1. Load
Load the Iris dataset using Scikit-Learn.
2. Prepare
Separate features X and target y.
3. Split
Create training and testing datasets.
4. Scale
Standardize numerical features.
5. Train
Fit the KNN classifier on training data.
6. Predict
Predict classes for unseen test samples.
7. Evaluate
Measure accuracy and classification metrics.
8. Deploy Logic
Use the trained model and scaler to classify new flowers.
16. Expected Result
You should typically see approximately 96–100% accuracy on the Iris dataset with KNN using this particular split and configuration. Exact results can vary when the train/test split, random seed, preprocessing, or model parameters change.
A high accuracy on this small, well-known dataset does not mean that every real-world classification problem will be equally easy. Real datasets can contain missing values, noise, class imbalance, outliers, irrelevant features, and much more complex relationships.
17. Common Mistakes
- Training the model before splitting the dataset.
- Fitting the scaler on the entire dataset.
- Using
fit_transform()on the test data. - Forgetting to scale new observations before prediction.
- Mixing up
Xandy. - Evaluating predictions against the wrong target values.
- Assuming high accuracy on Iris means every classification problem will be easy.
- Changing model parameters repeatedly based on the test set without a proper validation strategy.
18. Key Points
- The Iris dataset contains 150 samples, 4 features, and 3 classes.
Xcontains the flower measurements andycontains the target classes.train_test_split()separates training and testing data.StandardScalerstandardizes the numerical features.- KNN classifies samples based on nearby training observations.
n_neighborscontrols how many neighbors KNN considers.fit()trains the KNN model.predict()generates predictions for new observations.accuracy_score()measures overall classification accuracy.classification_report()provides detailed classification metrics.- The same fitted scaler must be used when processing future data.
Modify the project by testing different values of n_neighbors. Try 3, 5, 7, and 9. Train each model using the same training data and compare their test accuracy.
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score
for k in [3, 5, 7, 9]:
```
model = KNeighborsClassifier(n_neighbors=k)
# Train the model
model.fit(X_train, y_train)
# Make predictions
y_pred = model.predict(X_test)
# Calculate accuracy
accuracy = accuracy_score(y_test, y_pred)
print("K =", k, "| Accuracy =", accuracy)In Lecture 08, we will begin the next stage of the machine learning workflow: Model Evaluation. You will learn how to interpret evaluation metrics, confusion matrices, precision, recall, F1-score, and understand whether a model is actually performing well.
Linear Regression
Learn how to predict continuous values using Linear Regression, the most fundamental supervised learning algorithm. You will understand the mathematical idea behind a regression line, train a model with scikit-learn, evaluate its performance, and interpret what the model has learned.
- What linear regression is and how it works
- The difference between regression and classification
- How a regression line makes predictions
- How to train LinearRegression in sklearn
- How coefficients and intercepts are calculated and interpreted
- How to make predictions for new data
- How to evaluate regression models with MSE and R²
- How to understand underfitting and poor model performance
- How to avoid common regression mistakes
01. What Is Linear Regression?
Linear Regression is a supervised machine learning algorithm used to predict a continuous numerical value.
For example, we can use Linear Regression to predict:
- House price from house size
- Salary from years of experience
- Temperature from time of day
- Sales from advertising expenditure
- Student marks from study hours
- Fuel consumption from vehicle characteristics
The word linear means that the model attempts to describe the relationship between input variables and the target using a straight-line relationship.
Regression predicts numerical values such as 250000, 72.5, or 31.8.
Classification predicts categories such as "cat", "dog", "spam", or "not spam".
Our Iris project used classification because the target was a flower species. Linear Regression is different because its output is a continuous number.
02. The Basic Linear Regression Equation
For a single input feature, Linear Regression uses the equation:
y = mx + b
- y = predicted value
- x = input feature
- m = slope or coefficient
- b = intercept
In machine learning, this can also be written as:
Ε· = Ξ²β + Ξ²βx
Here, Ε· represents the predicted value, Ξ²β is the intercept, and Ξ²β is the coefficient of the feature.
x = 5 m = 10 b = 5 y = m * x + b print(y)
The result is 55 because:
y = (10 Γ 5) + 5 = 55
The trained Linear Regression model automatically learns suitable values for the coefficient and intercept from the training data.
03. Understanding the Coefficient
The coefficient tells us how much the predicted target changes when the input feature increases by one unit.
Suppose our model learns:
y = 8x + 20
The coefficient is 8.
This means that for every one-unit increase in x, the predicted value of y increases by approximately 8 units.
A positive coefficient means the prediction generally increases as the feature increases. A negative coefficient means the prediction generally decreases as the feature increases. A coefficient close to zero indicates a weak linear relationship, although its practical meaning depends on the scale and context of the feature.
04. Understanding the Intercept
The intercept is the predicted value of the target when all input features are zero.
For example:
y = 10x + 15
The intercept is 15.
When x = 0:
y = (10 Γ 0) + 15 = 15
In real-world datasets, the intercept may not always have a meaningful practical interpretation if a value of zero for the feature is outside the realistic range of the data.
05. How Linear Regression Learns
During training, Linear Regression searches for coefficients that produce predictions as close as possible to the actual target values.
The difference between an actual value and a predicted value is called the residual or error.
The model commonly uses the ordinary least squares approach, which finds parameters that minimize the sum of squared residuals.
Conceptually:
Residual = Actual Value β Predicted Value
Squaring the errors prevents positive and negative errors from cancelling each other out and gives larger errors more influence during optimization.
The model tries to find the line that provides the smallest overall squared prediction error for the training data.
06. Creating Training Data
Before training a regression model, we need two important pieces of information:
- X β the input features
- y β the target values we want to predict
For example, imagine that we want to predict a person's score based on the number of hours they studied.
import numpy as np X = np.array([[1], [2], [3], [4], [5]]) y = np.array([20, 35, 45, 55, 70]) print(X) print(y)
Here, X contains the number of study hours, while y contains the corresponding scores.
Scikit-learn expects the feature matrix X to normally have two dimensions: (number of samples, number of features). That is why we use [[1], [2], [3]] instead of [1, 2, 3] for X.
07. Importing LinearRegression
Scikit-learn provides Linear Regression through the sklearn.linear_model module.
from sklearn.linear_model import LinearRegression
We can then create a Linear Regression model:
model = LinearRegression()
At this point, the model has been created but it has not learned anything yet. Learning happens when we call fit().
08. Splitting the Dataset
Just like the Iris classification project, we should separate our data into training and testing sets.
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 )
The training data is used to learn the relationship between X and y. The testing data is kept separate so that we can evaluate how well the model performs on unseen examples.
If we evaluate the model using the same data it used for learning, we may get an overly optimistic view of its performance. A separate test set gives us a better estimate of how the model behaves on new data.
09. Training the Model
Once the data has been split, we can train the Linear Regression model using fit().
model = LinearRegression() model.fit(X_train, y_train)
During fit(), scikit-learn calculates the model parameters based on the training data.
After training, we can inspect the values learned by the model:
print("Coefficient:", model.coef_)
print("Intercept:", model.intercept_)10. Making Predictions
Once the model has been trained, we can use predict() to generate predictions for new data.
y_pred = model.predict(X_test) print(y_pred)
We can also predict the value for a completely new input.
new_data = np.array([[11]])
prediction = model.predict(new_data)
print("Prediction:", prediction[0])The model uses the equation it learned from the training data to calculate the predicted value.
11. Mean Squared Error (MSE)
Mean Squared Error, commonly called MSE, measures the average squared difference between actual values and predicted values.
The formula is:
MSE = (1/n) Γ Ξ£(actual β predicted)Β²
A lower MSE generally indicates that the predictions are closer to the actual values.
from sklearn.metrics import mean_squared_error
mse = mean_squared_error(y_test, y_pred)
print("MSE:", mse)MSE is expressed in squared units of the target. Therefore, its numerical value can sometimes be difficult to interpret directly. The scale of the target variable matters when deciding whether an MSE is large or small.
12. R² Score
R², or the coefficient of determination, measures how much of the variation in the target is explained by the regression model relative to a baseline that always predicts the mean target value.
A commonly used interpretation is:
- R² = 1.0 — perfect predictions on the evaluated data
- R² = 0 — equivalent to the mean-prediction baseline
- R² < 0 — worse than that baseline on the evaluated data
from sklearn.metrics import r2_score
score = r2_score(y_test, y_pred)
print("R2 Score:", score)R² should not be interpreted as a percentage of predictions that are correct. It describes explained variance relative to a mean-based baseline and must be interpreted in the context of the dataset and model.
13. Complete Linear Regression Example
Now let's put the complete workflow together: create data, split it, train the model, make predictions, and evaluate the results.
import numpy as np
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_squared_error, r2_score
X = np.array([[1],[2],[3],[4],[5],[6],[7],[8],[9],[10]])
y = np.array([15, 25, 35, 45, 55, 62, 70, 79, 88, 95])
X_train, X_test, y_train, y_test = train_test_split(
X,
y,
test_size=0.2,
random_state=42
)
model = LinearRegression()
model.fit(X_train, y_train)
y_pred = model.predict(X_test)
print("Coefficient:", model.coef_[0])
print("Intercept:", model.intercept_)
print("Predictions:", y_pred)
print("Actual:", y_test)
print("MSE:", mean_squared_error(y_test, y_pred))
print("R2 Score:", r2_score(y_test, y_pred))This example follows the standard supervised learning workflow:
Data β Split β Train β Predict β Evaluate
14. Interpreting the Model
Suppose the model produces:
Coefficient: 8.9 Intercept: 6.2 MSE: 4.5 R2 Score: 0.97
The coefficient tells us that an increase of one unit in X is associated with an increase of approximately 8.9 units in the predicted target, according to this fitted model.
The intercept is approximately 6.2, meaning the model predicts 6.2 when X is zero.
The R² value of 0.97 indicates that the fitted model explains a large proportion of the target variation relative to the mean baseline on this evaluated dataset. It does not mean that every prediction is exactly 97% correct.
15. Multiple Linear Regression
Linear Regression is not limited to one feature. We can use multiple features to predict a single continuous target.
For example, house price could depend on:
- House size
- Number of bedrooms
- Number of bathrooms
- Age of the house
The equation becomes:
y = Ξ²β + Ξ²βxβ + Ξ²βxβ + Ξ²βxβ + ... + Ξ²βxβ
import numpy as np
from sklearn.linear_model import LinearRegression
X = np.array([
[1000, 2],
[1500, 3],
[2000, 3],
[2500, 4],
[3000, 5]
])
y = np.array([150000, 220000, 300000, 380000, 450000])
model = LinearRegression()
model.fit(X, y)
print("Coefficients:", model.coef_)
print("Intercept:", model.intercept_)
new_house = np.array([[2200, 3]])
prediction = model.predict(new_house)
print("Predicted price:", prediction[0])Here, the first feature represents house size and the second feature represents the number of bedrooms. The model learns a separate coefficient for each feature.
16. Visualizing a Regression Line
For a dataset with one feature, visualization is an excellent way to understand Linear Regression. We can plot the original data points and then draw the line learned by the model.
import matplotlib.pyplot as plt
plt.scatter(X, y, label="Actual Data")
plt.plot(X, model.predict(X), label="Regression Line")
plt.xlabel("X")
plt.ylabel("y")
plt.title("Linear Regression")
plt.legend()
plt.show()The dots represent the observed training examples, while the line represents the predictions made by the model.
If the data points roughly follow a straight-line pattern, Linear Regression may be a useful starting model. If the relationship is strongly curved or much more complex, a simple linear model may not capture the pattern adequately.
17. Common Linear Regression Mistakes
Beginners often make several mistakes when building regression models.
- Using classification metrics: Accuracy is generally not the appropriate primary metric for continuous regression targets.
- Testing on training data: This can give an overly optimistic evaluation.
- Ignoring data leakage: Information from the test set should not influence training or preprocessing decisions.
- Assuming high R² means a perfect model: R² must be interpreted together with the prediction errors and application context.
- Ignoring outliers: Extreme observations can strongly influence ordinary least squares regression.
- Assuming correlation proves causation: A fitted relationship does not by itself establish that one variable causes another.
- Using linear regression for every problem: Some relationships are not well represented by a straight-line model.
Always separate the training and testing data before making data-driven preprocessing decisions. The test set should behave like unseen future data during model development.
18. Key Points to Remember
Regression
Predicts continuous numerical values rather than discrete classes.
Coefficient
Describes how the prediction changes with a one-unit change in a feature, holding other features constant in a multiple regression model.
Intercept
The model's predicted target when all input features are zero.
MSE
Measures the average squared prediction error. Lower values are generally better.
R²
Measures performance relative to a mean-based baseline.
fit()
Trains the Linear Regression model using the training data.
predict()
Uses the learned model to generate predictions for new feature values.
Test Data
Provides an independent evaluation of how the model performs on unseen examples.
- MSE — average squared difference between actual and predicted values. Lower is better when comparing models on the same target scale.
- R² Score — measures performance relative to predicting the mean target value. 1.0 represents a perfect fit on the evaluated data, while negative values are possible.
Create a dataset where X = house sizes (500–3000 sqft) and y = prices.
Train a LinearRegression model and perform the following tasks:
- Split the dataset into training and testing sets.
- Train the Linear Regression model.
- Print the coefficient and intercept.
- Predict the price of a 2000 sqft house.
- Calculate the MSE.
- Calculate the R² score.
- Try predicting prices for 1000, 1500, 2500, and 3000 sqft houses.
- Explain what the coefficient means in the context of house prices.
Use Matplotlib to create a scatter plot of the house sizes and prices. Then draw the regression line on the same graph. Look at the points and decide whether a straight-line relationship appears reasonable.
In the next lecture, we will explore another important supervised learning algorithm: Logistic Regression. You will learn how regression can also be used for classification problems by predicting probabilities and class labels.