CSPT lab programs
Required softwares
1. Python 3.10 or any other
2. Git
3. Visual studio code
4. Jupyter
Lab 1
Install these all libraries using cmd or notebook
pip install numpy pandas matplotlib seaborn scikit-learn notebook jupyter
Run this code to verify installation
if you want to run this through jupyter then run this code
pip install notebook
jupyter notebook
Program 1
Expected Output
Welcome to Machine Learning Lab
Numbers Square
0 10 100
1 20 400
2 30 900
3 40 1600
4 50 2500
pip list
to check libraries list from cmd or jupyter notebook
Lab 2: Load and explore the Iris dataset using Pandas.
Expected output:
==================================================
First Five Rows
==================================================
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) \
0 5.1 3.5 1.4 0.2
1 4.9 3.0 1.4 0.2
2 4.7 3.2 1.3 0.2
3 4.6 3.1 1.5 0.2
4 5.0 3.6 1.4 0.2
species
0 Setosa
1 Setosa
2 Setosa
3 Setosa
4 Setosa
Dataset Shape
(150, 5)
Dataset Information
<class 'pandas.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 sepal length (cm) 150 non-null float64
1 sepal width (cm) 150 non-null float64
2 petal length (cm) 150 non-null float64
3 petal width (cm) 150 non-null float64
4 species 150 non-null str
dtypes: float64(4), str(1)
memory usage: 7.2 KB
None
Summary Statistics
sepal length (cm) sepal width (cm) petal length (cm) \
count 150.000000 150.000000 150.000000
mean 5.843333 3.057333 3.758000
std 0.828066 0.435866 1.765298
min 4.300000 2.000000 1.000000
25% 5.100000 2.800000 1.600000
50% 5.800000 3.000000 4.350000
75% 6.400000 3.300000 5.100000
max 7.900000 4.400000 6.900000
petal width (cm)
count 150.000000
mean 1.199333
std 0.762238
min 0.100000
25% 0.300000
50% 1.300000
75% 1.800000
max 2.500000
Missing Values
sepal length (cm) 0
sepal width (cm) 0
petal length (cm) 0
petal width (cm) 0
species 0
dtype: int64
Species Distribution
species
Setosa 50
Versicolor 50
Virginica 50
Name: count, dtype: int64
Random Sample
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) \
74 6.4 2.9 4.3 1.3
121 5.6 2.8 4.9 2.0
58 6.6 2.9 4.6 1.3
73 6.1 2.8 4.7 1.2
87 6.3 2.3 4.4 1.3
species
74 Versicolor
121 Virginica
58 Versicolor
73 Versicolor
87 Versicolor
Dataset Description
.. _iris_dataset:
Iris plants dataset
--------------------
**Data Set Characteristics:**
:Number of Instances: 150 (50 in each of three classes)
:Number of Attributes: 4 numeric, predictive attributes and the class
:Attribute Information:
- sepal length in cm
- sepal width in cm
- petal length in cm
- petal width in cm
- class:
- Iris-Setosa
- Iris-Versicolour
- Iris-Virginica
:Summary Statistics:
============== ==== ==== ======= ===== ====================
Min Max Mean SD Class Correlation
============== ==== ==== ======= ===== ====================
sepal length: 4.3 7.9 5.84 0.83 0.7826
sepal width: 2.0 4.4 3.05 0.43 -0.4194
petal length: 1.0 6.9 3.76 1.76 0.9490 (high!)
petal width: 0.1 2.5 1.20 0.76 0.9565 (high!)
============== ==== ==== ======= ===== ====================
:Missing Attribute Values: None
:Class Distribution: 33.3% for each of 3 classes.
:Creator: R.A. Fisher
:Donor: Michael Marshall (MARSHALL%PLU@io.arc.nasa.gov)
:Date: July, 1988
The famous Iris database, first used by Sir R.A. Fisher. The dataset is taken
from Fisher's paper. Note that it's the same as in R, but not as in the UCI
Machine Learning Repository, which has two wrong data points.
This is perhaps the best known database to be found in the
pattern recognition literature. Fisher's paper is a classic in the field and
is referenced frequently to this day. (See Duda & Hart, for example.) The
data set contains 3 classes of 50 instances each, where each class refers to a
type of iris plant. One class is linearly separable from the other 2; the
latter are NOT linearly separable from each other.
.. dropdown:: References
- Fisher, R.A. "The use of multiple measurements in taxonomic problems"
Annual Eugenics, 7, Part II, 179-188 (1936); also in "Contributions to
Mathematical Statistics" (John Wiley, NY, 1950).
- Duda, R.O., & Hart, P.E. (1973) Pattern Classification and Scene Analysis.
(Q327.D83) John Wiley & Sons. ISBN 0-471-22361-1. See page 218.
- Dasarathy, B.V. (1980) "Nosing Around the Neighborhood: A New System
Structure and Classification Rule for Recognition in Partially Exposed
Environments". IEEE Transactions on Pattern Analysis and Machine
Intelligence, Vol. PAMI-2, No. 1, 67-71.
- Gates, G.W. (1972) "The Reduced Nearest Neighbor Rule". IEEE Transactions
on Information Theory, May 1972, 431-433.
- See also: 1988 MLC Proceedings, 54-64. Cheeseman et al"s AUTOCLASS II
conceptual clustering system finds 3 classes in the data.
- Many, many more ...
Lab 3: Exploratory Data Analysis (EDA) on the Iris Dataset
Aim
To perform Exploratory Data Analysis (EDA) on the Iris dataset using Pandas, Matplotlib, and Seaborn.
Objectives
After completing this lab, students will be able to:
- Load the Iris dataset.
- Generate summary statistics.
- Check missing values.
- Analyze feature correlations.
- Create pair plots.
- Plot histograms.
- Create box plots.
- Interpret patterns and relationships in the dataset.
Code:
Expected output:
==================================================
First Five Rows
==================================================
sepal length (cm) sepal width (cm) petal length (cm) petal width (cm) \
0 5.1 3.5 1.4 0.2
1 4.9 3.0 1.4 0.2
2 4.7 3.2 1.3 0.2
3 4.6 3.1 1.5 0.2
4 5.0 3.6 1.4 0.2
species
0 setosa
1 setosa
2 setosa
3 setosa
4 setosa
Dataset Shape
(150, 5)
Dataset Information
<class 'pandas.DataFrame'>
RangeIndex: 150 entries, 0 to 149
Data columns (total 5 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 sepal length (cm) 150 non-null float64
1 sepal width (cm) 150 non-null float64
2 petal length (cm) 150 non-null float64
3 petal width (cm) 150 non-null float64
4 species 150 non-null category
dtypes: category(1), float64(4)
memory usage: 5.0 KB
None
Summary Statistics
sepal length (cm) sepal width (cm) petal length (cm) \
count 150.000000 150.000000 150.000000
mean 5.843333 3.057333 3.758000
std 0.828066 0.435866 1.765298
min 4.300000 2.000000 1.000000
25% 5.100000 2.800000 1.600000
50% 5.800000 3.000000 4.350000
75% 6.400000 3.300000 5.100000
max 7.900000 4.400000 6.900000
petal width (cm)
count 150.000000
mean 1.199333
std 0.762238
min 0.100000
25% 0.300000
50% 1.300000
75% 1.800000
max 2.500000
Missing Values
sepal length (cm) 0
sepal width (cm) 0
petal length (cm) 0
petal width (cm) 0
species 0
dtype: int64
Species Count
species
setosa 50
versicolor 50
virginica 50
Name: count, dtype: int64
Lab 4: Apply Preprocessing Using Scikit-learn Pipelines
Aim
To learn how to preprocess data using Scikit-learn Pipelines by handling missing values, scaling numerical features, encoding categorical features, and preparing the dataset for machine learning.
Objectives
After completing this lab, students will be able to:
- Understand the importance of preprocessing.
- Create preprocessing pipelines using Scikit-learn.
- Handle missing values.
- Scale numerical features.
- Encode categorical variables.
- Split the dataset into training and testing sets.
- Train a machine learning model using a pipeline.
Code:
Expected output:
================================================== Model Accuracy ================================================== 1.0 Predicted Flower: setosa
Lab 5: Exploratory Data Analysis (EDA) on the California Housing Dataset
Aim
To perform Exploratory Data Analysis (EDA) on the California Housing dataset using Pandas, Matplotlib, and Seaborn to understand feature distributions, relationships, and correlations.
Objectives
After completing this lab, students will be able to:
- Load the California Housing dataset.
- Explore the dataset structure.
- Generate summary statistics.
- Check for missing values.
- Analyze feature correlations.
- Create histograms and box plots.
- Visualize feature relationships.
- Draw conclusions from the dataset.
Code:
Expected output:
==================================================
First Five Rows
==================================================
MedInc HouseAge AveRooms AveBedrms Population AveOccup Latitude \
0 8.3252 41.0 6.984127 1.023810 322.0 2.555556 37.88
1 8.3014 21.0 6.238137 0.971880 2401.0 2.109842 37.86
2 7.2574 52.0 8.288136 1.073446 496.0 2.802260 37.85
3 5.6431 52.0 5.817352 1.073059 558.0 2.547945 37.85
4 3.8462 52.0 6.281853 1.081081 565.0 2.181467 37.85
Longitude HouseValue
0 -122.23 4.526
1 -122.22 3.585
2 -122.24 3.521
3 -122.25 3.413
4 -122.25 3.422
Dataset Shape
(20640, 9)
Dataset Information
<class 'pandas.DataFrame'>
RangeIndex: 20640 entries, 0 to 20639
Data columns (total 9 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 MedInc 20640 non-null float64
1 HouseAge 20640 non-null float64
2 AveRooms 20640 non-null float64
3 AveBedrms 20640 non-null float64
4 Population 20640 non-null float64
5 AveOccup 20640 non-null float64
6 Latitude 20640 non-null float64
7 Longitude 20640 non-null float64
8 HouseValue 20640 non-null float64
dtypes: float64(9)
memory usage: 1.4 MB
None
Summary Statistics
MedInc HouseAge AveRooms AveBedrms Population \
count 20640.000000 20640.000000 20640.000000 20640.000000 20640.000000
mean 3.870671 28.639486 5.429000 1.096675 1425.476744
std 1.899822 12.585558 2.474173 0.473911 1132.462122
min 0.499900 1.000000 0.846154 0.333333 3.000000
25% 2.563400 18.000000 4.440716 1.006079 787.000000
50% 3.534800 29.000000 5.229129 1.048780 1166.000000
75% 4.743250 37.000000 6.052381 1.099526 1725.000000
max 15.000100 52.000000 141.909091 34.066667 35682.000000
AveOccup Latitude Longitude HouseValue
count 20640.000000 20640.000000 20640.000000 20640.000000
mean 3.070655 35.631861 -119.569704 2.068558
std 10.386050 2.135952 2.003532 1.153956
min 0.692308 32.540000 -124.350000 0.149990
25% 2.429741 33.930000 -121.800000 1.196000
50% 2.818116 34.260000 -118.490000 1.797000
75% 3.282261 37.710000 -118.010000 2.647250
max 1243.333333 41.950000 -114.310000 5.000010
Missing Values
MedInc 0
HouseAge 0
AveRooms 0
AveBedrms 0
Population 0
AveOccup 0
Latitude 0
Longitude 0
HouseValue 0
dtype: int64
Git hub link for this entire code
https://github.com/codingacharya/CSPT-programs.git
No comments:
Post a Comment