student notes / est. for the classroom

HTML, CSS, JavaScript, Python, data science, computer networks — written the way you'd explain it to a classmate, not a compiler.

Top Job & Internship Portals

Handpicked portals for fresher jobs, tech roles, and listings in Hyderabad

GFG

GeeksforGeeks

Tech & Software Roles

Visit →
INT

Internshala

Fresher Jobs & Internships

Visit →
GOOG

Google Careers

Global Google Openings

Visit →
APN

Apna Jobs

Local Jobs in Hyderabad

Visit →
INS

Instahyre

Tech Roles in Hyderabad

Visit →
NAUK

Naukri.com

Fresher Jobs in Hyderabad

Visit →
📢 Updated daily

Internship & Job Alerts

01

Latest notes

July 02, 2026

CSPT lab programs

 CSPT lab programs

 

Required softwares

1. Python 3.10 or any other

2. Git

3. Visual studio code

4. Jupyter

 

 Lab 1

Install these all libraries using cmd or notebook 

pip install numpy pandas matplotlib seaborn scikit-learn notebook jupyter

 

Run this code to verify installation  

 if you want to run this through jupyter then run this code

 

pip install notebook

jupyter notebook 

 

import sys
import numpy as np
import pandas as pd
import matplotlib
import seaborn as sns
import sklearn

print("=" * 50)
print("Machine Learning Environment Verification")
print("=" * 50)

print(f"Python Version      : {sys.version}")
print(f"NumPy Version       : {np.__version__}")
print(f"Pandas Version      : {pd.__version__}")
print(f"Matplotlib Version  : {matplotlib.__version__}")
print(f"Seaborn Version     : {sns.__version__}")
print(f"Scikit-learn Version: {sklearn.__version__}")

print("\nAll libraries installed successfully!")

 

 Program 1

print("Welcome to Machine Learning Lab")

import numpy as np
import pandas as pd

data = np.array([10, 20, 30, 40, 50])

df = pd.DataFrame({
    "Numbers": data,
    "Square": data**2
})

print(df)

 

 

Expected Output

Welcome to Machine Learning Lab

Numbers Square
0 10 100
1 20 400
2 30 900
3 40 1600
4 50 2500

 

 pip list

 to check libraries list from cmd or jupyter notebook

 

Lab 2: Load and explore the Iris dataset using Pandas. 

 

import pandas as pd
from sklearn.datasets import load_iris

# Load dataset
iris = load_iris()

# Create DataFrame
df = pd.DataFrame(
    iris.data,
    columns=iris.feature_names
)

# Add target labels
df["species"] = iris.target

# Convert numeric labels to names
df["species"] = df["species"].map({
    0: "Setosa",
    1: "Versicolor",
    2: "Virginica"
})

print("=" * 50)
print("First Five Rows")
print("=" * 50)
print(df.head())

print("\nDataset Shape")
print(df.shape)

print("\nDataset Information")
print(df.info())

print("\nSummary Statistics")
print(df.describe())

print("\nMissing Values")
print(df.isnull().sum())

print("\nSpecies Distribution")
print(df["species"].value_counts())

print("\nRandom Sample")
print(df.sample(5))

print("\nDataset Description")
print(iris.DESCR)

 

 Expected output:

==================================================
First Five Rows
==================================================
   sepal length (cm)  sepal width (cm)  petal length (cm)  petal width (cm)  \
0                5.1               3.5                1.4               0.2   
1                4.9               3.0                1.4               0.2   
2                4.7               3.2                1.3               0.2   
3                4.6               3.1                1.5               0.2   
4                5.0               3.6                1.4               0.2   

  species  
0  Setosa  
1  Setosa  
2  Setosa  
3  Setosa  
4  Setosa  

Dataset Shape
(150, 5)

Dataset Information
<class 'pandas.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
 #   Column             Non-Null Count  Dtype  
---  ------             --------------  -----  
 0   sepal length (cm)  150 non-null    float64
 1   sepal width (cm)   150 non-null    float64
 2   petal length (cm)  150 non-null    float64
 3   petal width (cm)   150 non-null    float64
 4   species            150 non-null    str    
dtypes: float64(4), str(1)
memory usage: 7.2 KB
None

Summary Statistics
       sepal length (cm)  sepal width (cm)  petal length (cm)  \
count         150.000000        150.000000         150.000000   
mean            5.843333          3.057333           3.758000   
std             0.828066          0.435866           1.765298   
min             4.300000          2.000000           1.000000   
25%             5.100000          2.800000           1.600000   
50%             5.800000          3.000000           4.350000   
75%             6.400000          3.300000           5.100000   
max             7.900000          4.400000           6.900000   

       petal width (cm)  
count        150.000000  
mean           1.199333  
std            0.762238  
min            0.100000  
25%            0.300000  
50%            1.300000  
75%            1.800000  
max            2.500000  

Missing Values
sepal length (cm)    0
sepal width (cm)     0
petal length (cm)    0
petal width (cm)     0
species              0
dtype: int64

Species Distribution
species
Setosa        50
Versicolor    50
Virginica     50
Name: count, dtype: int64

Random Sample
     sepal length (cm)  sepal width (cm)  petal length (cm)  petal width (cm)  \
74                 6.4               2.9                4.3               1.3   
121                5.6               2.8                4.9               2.0   
58                 6.6               2.9                4.6               1.3   
73                 6.1               2.8                4.7               1.2   
87                 6.3               2.3                4.4               1.3   

        species  
74   Versicolor  
121   Virginica  
58   Versicolor  
73   Versicolor  
87   Versicolor  

Dataset Description
.. _iris_dataset:

Iris plants dataset
--------------------

**Data Set Characteristics:**

:Number of Instances: 150 (50 in each of three classes)
:Number of Attributes: 4 numeric, predictive attributes and the class
:Attribute Information:
    - sepal length in cm
    - sepal width in cm
    - petal length in cm
    - petal width in cm
    - class:
            - Iris-Setosa
            - Iris-Versicolour
            - Iris-Virginica

:Summary Statistics:

============== ==== ==== ======= ===== ====================
                Min  Max   Mean    SD   Class Correlation
============== ==== ==== ======= ===== ====================
sepal length:   4.3  7.9   5.84   0.83    0.7826
sepal width:    2.0  4.4   3.05   0.43   -0.4194
petal length:   1.0  6.9   3.76   1.76    0.9490  (high!)
petal width:    0.1  2.5   1.20   0.76    0.9565  (high!)
============== ==== ==== ======= ===== ====================

:Missing Attribute Values: None
:Class Distribution: 33.3% for each of 3 classes.
:Creator: R.A. Fisher
:Donor: Michael Marshall (MARSHALL%PLU@io.arc.nasa.gov)
:Date: July, 1988

The famous Iris database, first used by Sir R.A. Fisher. The dataset is taken
from Fisher's paper. Note that it's the same as in R, but not as in the UCI
Machine Learning Repository, which has two wrong data points.

This is perhaps the best known database to be found in the
pattern recognition literature.  Fisher's paper is a classic in the field and
is referenced frequently to this day.  (See Duda & Hart, for example.)  The
data set contains 3 classes of 50 instances each, where each class refers to a
type of iris plant.  One class is linearly separable from the other 2; the
latter are NOT linearly separable from each other.

.. dropdown:: References

  - Fisher, R.A. "The use of multiple measurements in taxonomic problems"
    Annual Eugenics, 7, Part II, 179-188 (1936); also in "Contributions to
    Mathematical Statistics" (John Wiley, NY, 1950).
  - Duda, R.O., & Hart, P.E. (1973) Pattern Classification and Scene Analysis.
    (Q327.D83) John Wiley & Sons.  ISBN 0-471-22361-1.  See page 218.
  - Dasarathy, B.V. (1980) "Nosing Around the Neighborhood: A New System
    Structure and Classification Rule for Recognition in Partially Exposed
    Environments".  IEEE Transactions on Pattern Analysis and Machine
    Intelligence, Vol. PAMI-2, No. 1, 67-71.
  - Gates, G.W. (1972) "The Reduced Nearest Neighbor Rule".  IEEE Transactions
    on Information Theory, May 1972, 431-433.
  - See also: 1988 MLC Proceedings, 54-64.  Cheeseman et al"s AUTOCLASS II
    conceptual clustering system finds 3 classes in the data.
  - Many, many more ...

 

 

Lab 3: Exploratory Data Analysis (EDA) on the Iris Dataset

Aim

To perform Exploratory Data Analysis (EDA) on the Iris dataset using Pandas, Matplotlib, and Seaborn.


Objectives

After completing this lab, students will be able to:

  • Load the Iris dataset.
  • Generate summary statistics.
  • Check missing values.
  • Analyze feature correlations.
  • Create pair plots.
  • Plot histograms.
  • Create box plots.
  • Interpret patterns and relationships in the dataset.

 

 Code:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.datasets import load_iris

# Load dataset
iris = load_iris()

# Create DataFrame
df = pd.DataFrame(
    iris.data,
    columns=iris.feature_names
)

# Add species names
df["species"] = pd.Categorical.from_codes(
    iris.target,
    iris.target_names
)

# -----------------------------
# Basic Exploration
# -----------------------------
print("="*50)
print("First Five Rows")
print("="*50)
print(df.head())

print("\nDataset Shape")
print(df.shape)

print("\nDataset Information")
print(df.info())

print("\nSummary Statistics")
print(df.describe())

print("\nMissing Values")
print(df.isnull().sum())

print("\nSpecies Count")
print(df["species"].value_counts())

# -----------------------------
# Correlation Matrix
# -----------------------------
correlation = df.drop("species", axis=1).corr()

plt.figure(figsize=(8,6))
sns.heatmap(
    correlation,
    annot=True,
    cmap="coolwarm",
    linewidths=0.5
)
plt.title("Correlation Heatmap")
plt.show()

# -----------------------------
# Pair Plot
# -----------------------------
sns.pairplot(
    df,
    hue="species",
    diag_kind="hist"
)
plt.show()

# -----------------------------
# Histograms
# -----------------------------
df.hist(
    figsize=(10,8),
    bins=15
)
plt.suptitle("Feature Histograms")
plt.show()

# -----------------------------
# Box Plots
# -----------------------------
plt.figure(figsize=(12,6))

for i, column in enumerate(df.columns[:-1], 1):
    plt.subplot(2,2,i)
    sns.boxplot(y=df[column])
    plt.title(column)

plt.tight_layout()
plt.show()

# -----------------------------
# Species-wise Box Plot
# -----------------------------
plt.figure(figsize=(10,6))
sns.boxplot(
    x="species",
    y="petal length (cm)",
    data=df
)
plt.title("Petal Length by Species")
plt.show()

# -----------------------------
# Scatter Plot
# -----------------------------
plt.figure(figsize=(8,6))
sns.scatterplot(
    x="sepal length (cm)",
    y="petal length (cm)",
    hue="species",
    data=df
)
plt.title("Sepal Length vs Petal Length")
plt.show()

 

 Expected output:

 

Lab 4: Apply Preprocessing Using Scikit-learn Pipelines

 

Aim

To learn how to preprocess data using Scikit-learn Pipelines by handling missing values, scaling numerical features, encoding categorical features, and preparing the dataset for machine learning.


Objectives

After completing this lab, students will be able to:

  • Understand the importance of preprocessing.
  • Create preprocessing pipelines using Scikit-learn.
  • Handle missing values.
  • Scale numerical features.
  • Encode categorical variables.
  • Split the dataset into training and testing sets.
  • Train a machine learning model using a pipeline.

 

 

Code:

 

import pandas as pd

from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split

from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer

from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler

from sklearn.linear_model import LogisticRegression

from sklearn.metrics import accuracy_score

# -----------------------------
# Load Dataset
# -----------------------------

iris = load_iris()

df = pd.DataFrame(
    iris.data,
    columns=iris.feature_names
)

df["species"] = iris.target

# -----------------------------
# Create Missing Values
# -----------------------------

df.loc[5, "sepal length (cm)"] = None
df.loc[20, "petal width (cm)"] = None

# -----------------------------
# Split Features and Target
# -----------------------------

X = df.drop("species", axis=1)
y = df["species"]

X_train, X_test, y_train, y_test = train_test_split(
    X,
    y,
    test_size=0.20,
    random_state=42
)

# -----------------------------
# Pipeline
# -----------------------------

numeric_features = X.columns

numeric_pipeline = Pipeline([
    ("imputer", SimpleImputer(strategy="mean")),
    ("scaler", StandardScaler())
])

preprocessor = ColumnTransformer([
    ("num", numeric_pipeline, numeric_features)
])

pipeline = Pipeline([
    ("preprocessor", preprocessor),
    ("classifier", LogisticRegression(max_iter=200))
])

# -----------------------------
# Train
# -----------------------------

pipeline.fit(X_train, y_train)

# -----------------------------
# Predict
# -----------------------------

predictions = pipeline.predict(X_test)

accuracy = accuracy_score(y_test, predictions)

print("=" * 50)
print("Model Accuracy")
print("=" * 50)

print(accuracy)

# -----------------------------
# Predict New Flower
# -----------------------------

sample = pd.DataFrame({
    "sepal length (cm)": [5.2],
    "sepal width (cm)": [3.5],
    "petal length (cm)": [1.5],
    "petal width (cm)": [0.2]
})

result = pipeline.predict(sample)

print("\nPredicted Flower:", iris.target_names[result][0])

 

 Expected output:

==================================================
Model Accuracy
==================================================
1.0

Predicted Flower: setosa

 

 

Lab 5: Exploratory Data Analysis (EDA) on the California Housing Dataset

Aim

To perform Exploratory Data Analysis (EDA) on the California Housing dataset using Pandas, Matplotlib, and Seaborn to understand feature distributions, relationships, and correlations.


Objectives

After completing this lab, students will be able to:

  • Load the California Housing dataset.
  • Explore the dataset structure.
  • Generate summary statistics.
  • Check for missing values.
  • Analyze feature correlations.
  • Create histograms and box plots.
  • Visualize feature relationships.
  • Draw conclusions from the dataset.

 

 Code:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

from sklearn.datasets import fetch_california_housing

# -----------------------------
# Load Dataset
# -----------------------------

housing = fetch_california_housing()

df = pd.DataFrame(
    housing.data,
    columns=housing.feature_names
)

df["HouseValue"] = housing.target

# -----------------------------
# Basic Information
# -----------------------------

print("="*50)
print("First Five Rows")
print("="*50)

print(df.head())

print("\nDataset Shape")
print(df.shape)

print("\nDataset Information")
print(df.info())

print("\nSummary Statistics")
print(df.describe())

print("\nMissing Values")
print(df.isnull().sum())

# -----------------------------
# Correlation
# -----------------------------

correlation = df.corr()

plt.figure(figsize=(10,8))

sns.heatmap(
    correlation,
    annot=True,
    cmap="coolwarm"
)

plt.title("Correlation Heatmap")
plt.show()

# -----------------------------
# Histograms
# -----------------------------

df.hist(
    figsize=(15,10),
    bins=20
)

plt.suptitle("Feature Distributions")
plt.show()

# -----------------------------
# Boxplots
# -----------------------------

plt.figure(figsize=(15,10))

for i, column in enumerate(df.columns,1):

    plt.subplot(3,3,i)

    sns.boxplot(y=df[column])

    plt.title(column)

plt.tight_layout()

plt.show()

# -----------------------------
# Scatter Plot
# -----------------------------

plt.figure(figsize=(8,6))

sns.scatterplot(
    x="MedInc",
    y="HouseValue",
    data=df,
    alpha=0.4
)

plt.title("Median Income vs House Value")

plt.show()

# -----------------------------
# Pair Plot
# -----------------------------

selected = df[[
    "MedInc",
    "HouseAge",
    "AveRooms",
    "HouseValue"
]]

sns.pairplot(selected)

plt.show()

# -----------------------------
# House Value Distribution
# -----------------------------

plt.figure(figsize=(8,6))

sns.histplot(
    df["HouseValue"],
    bins=30,
    kde=True
)

plt.title("Distribution of House Values")

plt.show()

# -----------------------------
# Regression Plot
# -----------------------------

plt.figure(figsize=(8,6))

sns.regplot(
    x="MedInc",
    y="HouseValue",
    data=df,
    scatter_kws={"alpha":0.3}
)

plt.title("Median Income vs House Value")

plt.show()

 

 Expected output:

 

 

 Git hub link for this entire code

https://github.com/codingacharya/CSPT-programs.git 

 

 

 

 

No comments:

Post a Comment

02

Capstone resource hub

Codingacharya

Capstone Learning Resources, Notes & Project Hub

TCS NQT Questions
Read Notes
Machine Learning – ACE Theory
Read Notes
Machine Learning PPT
Read Notes
MachienLearning LAB
Read Notes
CSPT LAB programs
Read Notes
Time table and CSPT syllabus
Read Notes
Appreciations
Read Notes
ISTE life memberships
Read Notes
Artificial Intelligence & Analytics
Read Notes
Fullstack Web Dev
Read Notes
MERN Web Dev
Read Notes
Course Structure
Read Notes
Cloud Computing
Read Notes
90 Days ML Challenge
Read Notes
Advanced Analytics & Viz
Read Notes
Advanced Machine Learning
Read Notes
React JS
Read Notes
ML Chaitanya
Read Notes
Important Links
Read Notes
CSS Effects
Read Notes
RESUME
Read Notes
Bootstrap CSS
Read Notes
MongoDB
Read Notes
OWN Python Package
Read Notes
HTML Course
Read Notes
HTML Projects
Read Notes
GitHub Projects
Read Notes
Angular JS
Read Notes
Journals
Read Notes
NLP Notes
Read Notes
Videos
Read Notes
Data Analytics & Viz
Read Notes
Cloud Computing (Archive)
Read Notes
Open CV
Read Notes
jQuery
Read Notes
React JS (Archive)
Read Notes
Node JS
Read Notes
DAV Theory
Read Notes
DAV Lab
Read Notes
Big Data Notes
Read Notes
R-Programming
Read Notes
HADOOP Lab
Read Notes
GATE DA
Read Notes
JAVA Lab
Read Notes
Computer Networks
Read Notes
03

Live projects & profiles