student notes / est. for the classroom

HTML, CSS, JavaScript, Python, data science, computer networks — written the way you'd explain it to a classmate, not a compiler.

Top Job & Internship Portals

Handpicked portals for fresher jobs, tech roles, and listings in Hyderabad

GFG

GeeksforGeeks

Tech & Software Roles

Visit →
INT

Internshala

Fresher Jobs & Internships

Visit →
GOOG

Google Careers

Global Google Openings

Visit →
APN

Apna Jobs

Local Jobs in Hyderabad

Visit →
INS

Instahyre

Tech Roles in Hyderabad

Visit →
NAUK

Naukri.com

Fresher Jobs in Hyderabad

Visit →
📢 Updated daily

Internship & Job Alerts

01

Latest notes

November 27, 2025

Artificial Intelligence and Analytics CSE

 

Unit 1: Fundamentals of Artificial Intelligence

Topics:

  • Introduction to AI: History, Definitions, and Applications

  • Types of AI: Narrow AI, General AI, and Superintelligent AI

  • Intelligent Agents and Environments

  • Problem Solving and Search Techniques:

    • Uninformed Search: BFS, DFS

    • Informed Search: A*, Greedy Best-First

  • Introduction to Knowledge Representation and Reasoning

  • Basics of Machine Learning and AI vs ML vs DL

Practical:

  • Implement search algorithms in Python

  • Build a simple rule-based AI system


Unit 2: Machine Learning & Data Analytics

Topics:

  • Introduction to Machine Learning:

    • Supervised Learning: Regression, Classification

    • Unsupervised Learning: Clustering, Dimensionality Reduction

  • Feature Engineering and Preprocessing

  • Evaluation Metrics: Accuracy, Precision, Recall, F1-score, RMSE

  • Data Analytics Concepts:

    • Data Collection, Cleaning, and Visualization

    • Descriptive, Diagnostic, Predictive, and Prescriptive Analytics

  • Tools: Python libraries (NumPy, Pandas, Matplotlib, Seaborn, Scikit-learn)

Practical:

  • Build a regression and classification model

  • Analyze datasets to generate insights using Python


Unit 3: Deep Learning & AI Techniques

Topics:

  • Introduction to Deep Learning:

    • Neural Networks, Perceptron, Activation Functions

    • Forward and Backpropagation

  • Convolutional Neural Networks (CNN) for Image Analytics

  • Recurrent Neural Networks (RNN) & LSTM for Sequential Data

  • Natural Language Processing (NLP) basics

  • Reinforcement Learning fundamentals

  • Introduction to AI frameworks: TensorFlow, PyTorch

Practical:

  • Build a simple image classifier using CNN

  • Implement a text classification or sentiment analysis model


Unit 4: Advanced Analytics, AI Applications & Deployment

Topics:

  • Advanced Analytics: Predictive, Prescriptive, and Real-Time Analytics

  • Big Data Analytics (Hadoop, Spark)

  • AI in real-world applications: Healthcare, Finance, E-commerce, Robotics

  • AI Model Deployment and Monitoring

  • Ethics, Bias, and Explainable AI (XAI)

  • AI and Analytics trends: AutoML, Generative AI, Chatbots

Practical:

  • Develop an end-to-end AI project (e.g., sales prediction, chatbot, or recommendation system)

  • Deploy a trained AI model using Flask/Django or cloud platforms





Lab Program 1: Implementation of Uninformed Search Algorithms

Aim:
Implement Breadth First Search (BFS) and Depth First Search (DFS) for a given state-space problem.

Tasks:

  • Represent a graph using adjacency list

  • Implement BFS and DFS

  • Compare time and space complexity

Unit Covered: Unit 1 – Problem Solving & Search
Tools: Python


Lab Program 2: Implementation of Informed Search Algorithms

Aim:
Implement Greedy Best-First Search and A* algorithm.

Tasks:

  • Define heuristic function

  • Implement Greedy and A*

  • Compare optimality and efficiency

Use Case: Shortest path / puzzle problem
Unit Covered: Unit 1


Lab Program 3: Rule-Based Expert System

Aim:
Build a simple rule-based AI system.

Tasks:

  • Define facts and rules

  • Implement inference using IF-THEN rules

  • Example: Medical diagnosis / Career recommendation

Unit Covered: Unit 1 – Knowledge Representation
Tools: Python


Lab Program 4: Data Preprocessing & Feature Engineering

Aim:
Perform data preprocessing and feature engineering on a real dataset.

Tasks:

  • Handle missing values

  • Encode categorical variables

  • Normalize / standardize data

  • Visualize data distributions

Unit Covered: Unit 2 – Data Analytics
Tools: Pandas, NumPy, Matplotlib, Seaborn


Lab Program 5: Supervised Learning – Regression Model

Aim:
Build and evaluate a regression model.

Tasks:

  • Implement Linear Regression

  • Train-test split

  • Evaluate using RMSE and R²

Dataset: House price / sales prediction
Unit Covered: Unit 2
Tools: Scikit-learn


Lab Program 6: Supervised Learning – Classification Model

Aim:
Build a classification model.

Tasks:

  • Implement Logistic Regression / Decision Tree

  • Evaluate using Accuracy, Precision, Recall, F1-score

  • Confusion matrix visualization

Dataset: Spam detection / disease prediction
Unit Covered: Unit 2


Lab Program 7: Unsupervised Learning – Clustering

Aim:
Perform clustering on unlabeled data.

Tasks:

  • Implement K-Means clustering

  • Elbow method

  • Visualize clusters

Use Case: Customer segmentation
Unit Covered: Unit 2


Lab Program 8: Image Classification using CNN

Aim:
Build a Convolutional Neural Network for image classification.

Tasks:

  • Load image dataset

  • Build CNN architecture

  • Train and evaluate model

Dataset: MNIST / CIFAR-10
Unit Covered: Unit 3
Tools: TensorFlow / PyTorch


Lab Program 9: NLP – Text Classification / Sentiment Analysis

Aim:
Implement a text classification or sentiment analysis system.

Tasks:

  • Text preprocessing (tokenization, stopwords)

  • Build ML / DL model

  • Predict sentiment

Dataset: Movie reviews / tweets
Unit Covered: Unit 3


Lab Program 10: End-to-End AI Project with Deployment

Aim:
Develop and deploy an AI-powered application.

Options (choose one):

  • Sales prediction system

  • Recommendation system

  • AI chatbot

  • Fraud detection system

Tasks:

  • Data preprocessing

  • Model training

  • API deployment using Flask/Django

  • Basic monitoring & ethics discussion

Unit Covered: Unit 4
Tools: Flask/Django, Python, ML/DL libraries




LAB 1

✅ Streamlit Code

Save this as app.py and run using:

streamlit run app.py
import streamlit as st from collections import deque st.set_page_config(page_title="Uninformed Search Algorithms", layout="wide") st.title("🔍 Lab Program 1: Uninformed Search Algorithms") st.subheader("Breadth First Search (BFS) and Depth First Search (DFS)") # ----------------------------- # Graph Input Section # ----------------------------- st.sidebar.header("Graph Input") num_nodes = st.sidebar.number_input("Number of Nodes", min_value=1, max_value=20, value=5) st.sidebar.write("Enter edges (format: A B means A -> B)") edges_input = st.sidebar.text_area("Edges (one per line)", """0 1 0 2 1 3 2 4""") start_node = st.sidebar.text_input("Start Node", "0") # ----------------------------- # Create Graph (Adjacency List) # ----------------------------- def create_graph(edges_text): graph = {} lines = edges_text.strip().split("\n") for line in lines: if line.strip(): u, v = line.split() if u not in graph: graph[u] = [] if v not in graph: graph[v] = [] graph[u].append(v) return graph graph = create_graph(edges_input) st.subheader("📌 Adjacency List Representation") st.json(graph) # ----------------------------- # BFS Implementation # ----------------------------- def bfs(graph, start): visited = set() queue = deque([start]) traversal = [] while queue: node = queue.popleft() if node not in visited: visited.add(node) traversal.append(node) queue.extend(graph[node]) return traversal # ----------------------------- # DFS Implementation # ----------------------------- def dfs(graph, start, visited=None, traversal=None): if visited is None: visited = set() traversal = [] visited.add(start) traversal.append(start) for neighbor in graph[start]: if neighbor not in visited: dfs(graph, neighbor, visited, traversal) return traversal # ----------------------------- # Execute Algorithms # ----------------------------- if st.button("Run BFS & DFS"): if start_node not in graph: st.error("Start node not found in graph!") else: col1, col2 = st.columns(2) with col1: st.subheader("🔵 BFS Traversal") bfs_result = bfs(graph, start_node) st.success(" → ".join(bfs_result)) with col2: st.subheader("🟢 DFS Traversal") dfs_result = dfs(graph, start_node) st.success(" → ".join(dfs_result)) # ----------------------------- # Complexity Comparison # ----------------------------- st.subheader("⏱ Time & Space Complexity Comparison") st.markdown(""" ### 🔵 Breadth First Search (BFS) - **Time Complexity:** O(V + E) - **Space Complexity:** O(V) - Uses Queue (FIFO) - Explores level by level ### 🟢 Depth First Search (DFS) - **Time Complexity:** O(V + E) - **Space Complexity:** O(V) - Uses Recursion / Stack - Explores depth first --- ### 📊 Comparison Summary | Algorithm | Time Complexity | Space Complexity | Data Structure Used | |-----------|----------------|------------------|---------------------| | BFS | O(V + E) | O(V) | Queue | | DFS | O(V + E) | O(V) | Stack / Recursion | Where: - V = Number of vertices - E = Number of edges """) st.info("Both BFS and DFS have the same time complexity, but BFS generally consumes more memory for wide graphs.")

📌 What This Program Does

✔ Accepts graph input
✔ Converts it to adjacency list
✔ Runs BFS
✔ Runs DFS
✔ Displays traversal order
✔ Shows complexity comparison


LAB 2

Here is a complete Streamlit application for:

🔍 Lab Program 2: Informed Search Algorithms

Greedy Best-First Search & A* Algorithm

This program includes:

  • ✅ Heuristic function definition

  • ✅ Greedy Best-First Search implementation

  • ✅ A* algorithm implementation

  • ✅ Optimality & efficiency comparison

  • ✅ Interactive visualization (cost + path)


▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import heapq st.set_page_config(page_title="Informed Search Algorithms", layout="wide") st.title("🔍 Lab Program 2: Informed Search Algorithms") st.subheader("Greedy Best-First Search and A* Algorithm") # --------------------------------------- # Sidebar Inputs # --------------------------------------- st.sidebar.header("Graph Input") edges_input = st.sidebar.text_area( "Enter edges with cost (format: A B 1)", """A B 1 A C 4 B D 2 C D 1 D E 3""" ) heuristic_input = st.sidebar.text_area( "Enter heuristic values (format: Node Value)", """A 7 B 6 C 2 D 1 E 0""" ) start = st.sidebar.text_input("Start Node", "A") goal = st.sidebar.text_input("Goal Node", "E") # --------------------------------------- # Graph Creation # --------------------------------------- def create_graph(edges_text): graph = {} for line in edges_text.strip().split("\n"): u, v, cost = line.split() cost = int(cost) if u not in graph: graph[u] = [] if v not in graph: graph[v] = [] graph[u].append((v, cost)) return graph def create_heuristic(h_text): h = {} for line in h_text.strip().split("\n"): node, value = line.split() h[node] = int(value) return h graph = create_graph(edges_input) heuristic = create_heuristic(heuristic_input) st.subheader("📌 Graph Representation (Adjacency List)") st.json(graph) st.subheader("📌 Heuristic Function h(n)") st.json(heuristic) # --------------------------------------- # Greedy Best-First Search # --------------------------------------- def greedy_best_first(graph, start, goal, h): visited = set() pq = [(h[start], start, 0, [start])] # (priority, node, cost, path) while pq: _, node, cost, path = heapq.heappop(pq) if node == goal: return path, cost if node not in visited: visited.add(node) for neighbor, edge_cost in graph[node]: if neighbor not in visited: heapq.heappush(pq, (h[neighbor], neighbor, cost + edge_cost, path + [neighbor])) return None, None # --------------------------------------- # A* Algorithm # --------------------------------------- def a_star(graph, start, goal, h): pq = [(h[start], start, 0, [start])] # (f(n), node, g(n), path) visited = {} while pq: f, node, g, path = heapq.heappop(pq) if node == goal: return path, g if node not in visited or g < visited[node]: visited[node] = g for neighbor, cost in graph[node]: g_new = g + cost f_new = g_new + h[neighbor] heapq.heappush(pq, (f_new, neighbor, g_new, path + [neighbor])) return None, None # --------------------------------------- # Execute # --------------------------------------- if st.button("Run Greedy & A*"): if start not in graph or goal not in graph: st.error("Invalid Start or Goal Node") else: col1, col2 = st.columns(2) with col1: st.subheader("🟢 Greedy Best-First Search") path_g, cost_g = greedy_best_first(graph, start, goal, heuristic) if path_g: st.success("Path: " + " → ".join(path_g)) st.info(f"Total Cost: {cost_g}") else: st.error("No Path Found") with col2: st.subheader("🔵 A* Search") path_a, cost_a = a_star(graph, start, goal, heuristic) if path_a: st.success("Path: " + " → ".join(path_a)) st.info(f"Total Cost: {cost_a}") else: st.error("No Path Found") # --------------------------------------- # Comparison Section # --------------------------------------- st.subheader("📊 Optimality & Efficiency Comparison") st.markdown(""" ### 🟢 Greedy Best-First Search - Uses: **f(n) = h(n)** - Focuses only on heuristic - Faster in many cases - ❌ Not always optimal - ❌ Can get stuck in local minima ### 🔵 A* Algorithm - Uses: **f(n) = g(n) + h(n)** - Considers both actual cost and heuristic - ✅ Complete - ✅ Optimal (if heuristic is admissible) - Slightly higher memory usage --- ### 📌 Summary Table | Algorithm | Evaluation Function | Optimal | Complete | Speed | |------------|--------------------|----------|------------|--------| | Greedy | f(n) = h(n) | ❌ No | ❌ Not Always | Fast | | A* | f(n) = g(n)+h(n) | ✅ Yes* | ✅ Yes | Moderate | \* A* is optimal if heuristic is admissible (never overestimates). """) st.info("A* generally gives better paths, while Greedy is faster but may produce suboptimal solutions.")

📘 Theory (For Record Submission)

🔹 Heuristic Function

A heuristic function h(n) estimates the cost from node n to the goal node.

🔹 Greedy Best-First Search

  • Chooses node with lowest heuristic value

  • Does not consider actual path cost

  • May not give optimal solution

🔹 A* Algorithm

  • Uses:

    f(n)=g(n)+h(n)f(n) = g(n) + h(n)

    where:

    • g(n) = actual cost from start

    • h(n) = estimated cost to goal

  • Optimal if heuristic is admissible


LAB 3

▶️ How to Run

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st st.set_page_config(page_title="Rule-Based Expert System", layout="wide") st.title("🧠 Lab Program 3: Rule-Based Expert System") st.subheader("IF–THEN Rule Based Inference System") # ------------------------------------------------ # Select System # ------------------------------------------------ system_type = st.selectbox( "Choose Expert System Type:", ["Medical Diagnosis", "Career Recommendation"] ) # ------------------------------------------------ # Forward Chaining Engine # ------------------------------------------------ def forward_chaining(facts, rules): inferred = set(facts) applied_rules = [] changed = True while changed: changed = False for rule in rules: if rule["if"].issubset(inferred) and rule["then"] not in inferred: inferred.add(rule["then"]) applied_rules.append(rule) changed = True return inferred, applied_rules # ================================================= # 1️⃣ MEDICAL DIAGNOSIS SYSTEM # ================================================= if system_type == "Medical Diagnosis": st.header("🩺 Medical Diagnosis Expert System") st.write("Select symptoms:") fever = st.checkbox("Fever") cough = st.checkbox("Cough") headache = st.checkbox("Headache") fatigue = st.checkbox("Fatigue") sore_throat = st.checkbox("Sore Throat") # Facts facts = set() if fever: facts.add("fever") if cough: facts.add("cough") if headache: facts.add("headache") if fatigue: facts.add("fatigue") if sore_throat: facts.add("sore_throat") # Rules rules = [ {"if": {"fever", "cough"}, "then": "flu"}, {"if": {"fever", "headache"}, "then": "migraine"}, {"if": {"cough", "sore_throat"}, "then": "common_cold"}, {"if": {"fatigue", "fever"}, "then": "viral_infection"}, ] if st.button("Diagnose"): inferred, applied = forward_chaining(facts, rules) st.subheader("🔎 Inference Process") st.write("Initial Facts:", facts) for r in applied: st.write(f"Applied Rule: IF {r['if']} THEN {r['then']}") diagnosis = inferred - facts if diagnosis: st.success("Possible Diagnosis: " + ", ".join(diagnosis)) else: st.warning("No specific diagnosis found.") # ================================================= # 2️⃣ CAREER RECOMMENDATION SYSTEM # ================================================= else: st.header("🎓 Career Recommendation Expert System") st.write("Select your interests/skills:") math = st.checkbox("Strong in Mathematics") coding = st.checkbox("Interested in Coding") biology = st.checkbox("Interested in Biology") creativity = st.checkbox("Creative Thinking") communication = st.checkbox("Good Communication Skills") # Facts facts = set() if math: facts.add("math") if coding: facts.add("coding") if biology: facts.add("biology") if creativity: facts.add("creativity") if communication: facts.add("communication") # Rules rules = [ {"if": {"math", "coding"}, "then": "Software Engineer"}, {"if": {"biology", "math"}, "then": "Doctor"}, {"if": {"creativity"}, "then": "Designer"}, {"if": {"communication", "creativity"}, "then": "Marketing Manager"}, ] if st.button("Recommend Career"): inferred, applied = forward_chaining(facts, rules) st.subheader("🔎 Inference Process") st.write("Initial Facts:", facts) for r in applied: st.write(f"Applied Rule: IF {r['if']} THEN {r['then']}") recommendation = inferred - facts if recommendation: st.success("Recommended Career: " + ", ".join(recommendation)) else: st.warning("No specific recommendation found.") # ------------------------------------------------ # Theory & Comparison Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 What is a Rule-Based Expert System? A rule-based expert system is an AI system that uses predefined **IF–THEN rules** to make decisions. ### 🔹 Components: 1. Knowledge Base (Facts + Rules) 2. Inference Engine 3. User Interface ### 🔹 Inference Method Used: Forward Chaining (Data-Driven) --- ### 📊 Advantages: - Simple to implement - Transparent reasoning - Easy to modify rules ### ❌ Limitations: - Cannot learn automatically - Rule explosion problem - Limited to predefined knowledge """) st.info("This system uses Forward Chaining: It starts with known facts and applies rules to infer new conclusions.")

📘 Record Explanation (For Submission)

🔹 Facts

Facts represent known information given by the user.

Example:

fever = True cough = True

🔹 Rules

Rules are in IF–THEN format:

IF fever AND cough THEN flu

🔹 Inference Engine

Uses Forward Chaining:

  • Start with known facts

  • Apply matching rules

  • Infer new facts

  • Continue until no new facts are generated


🎯 Output Example (Medical)

Input:

  • Fever ✔

  • Cough ✔

Inference:

IF fever AND cough THEN flu

Output:

Possible Diagnosis: flu




LAB 4

Here is a complete Streamlit Lab Program 4 implementation for:

📊 Lab Program 4: Data Preprocessing & Feature Engineering

Real Dataset Processing (Upload Your Own CSV)

This app allows you to:

  • ✅ Upload a real dataset (CSV)

  • ✅ Handle missing values

  • ✅ Encode categorical variables

  • ✅ Normalize / Standardize data

  • ✅ Visualize distributions


▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.preprocessing import LabelEncoder, MinMaxScaler, StandardScaler st.set_page_config(page_title="Data Preprocessing & Feature Engineering", layout="wide") st.title("📊 Lab Program 4: Data Preprocessing & Feature Engineering") # ------------------------------------------------ # Upload Dataset # ------------------------------------------------ st.sidebar.header("Upload Dataset") file = st.sidebar.file_uploader("Upload CSV file", type=["csv"]) if file is not None: df = pd.read_csv(file) st.subheader("📌 Original Dataset") st.write(df.head()) st.write("Shape:", df.shape) # ------------------------------------------------ # 1️⃣ Handle Missing Values # ------------------------------------------------ st.subheader("1️⃣ Handle Missing Values") missing = df.isnull().sum() st.write("Missing Values per Column:") st.write(missing) missing_option = st.selectbox( "Choose Missing Value Strategy:", ["None", "Drop Rows", "Fill with Mean (Numeric)", "Fill with Median (Numeric)", "Fill with Mode"] ) if missing_option == "Drop Rows": df = df.dropna() elif missing_option == "Fill with Mean (Numeric)": for col in df.select_dtypes(include=np.number).columns: df[col].fillna(df[col].mean(), inplace=True) elif missing_option == "Fill with Median (Numeric)": for col in df.select_dtypes(include=np.number).columns: df[col].fillna(df[col].median(), inplace=True) elif missing_option == "Fill with Mode": for col in df.columns: df[col].fillna(df[col].mode()[0], inplace=True) st.success("Missing values handled.") st.write(df.head()) # ------------------------------------------------ # 2️⃣ Encode Categorical Variables # ------------------------------------------------ st.subheader("2️⃣ Encode Categorical Variables") categorical_cols = df.select_dtypes(include="object").columns.tolist() st.write("Categorical Columns:", categorical_cols) encoding_option = st.selectbox( "Choose Encoding Method:", ["None", "Label Encoding", "One Hot Encoding"] ) if encoding_option == "Label Encoding": le = LabelEncoder() for col in categorical_cols: df[col] = le.fit_transform(df[col]) elif encoding_option == "One Hot Encoding": df = pd.get_dummies(df, columns=categorical_cols) st.success("Encoding Completed.") st.write(df.head()) # ------------------------------------------------ # 3️⃣ Normalization / Standardization # ------------------------------------------------ st.subheader("3️⃣ Normalize / Standardize Data") numeric_cols = df.select_dtypes(include=np.number).columns.tolist() st.write("Numeric Columns:", numeric_cols) scaling_option = st.selectbox( "Choose Scaling Method:", ["None", "Min-Max Normalization", "Standardization (Z-score)"] ) if scaling_option == "Min-Max Normalization": scaler = MinMaxScaler() df[numeric_cols] = scaler.fit_transform(df[numeric_cols]) elif scaling_option == "Standardization (Z-score)": scaler = StandardScaler() df[numeric_cols] = scaler.fit_transform(df[numeric_cols]) st.success("Scaling Applied.") st.write(df.head()) # ------------------------------------------------ # 4️⃣ Data Visualization # ------------------------------------------------ st.subheader("4️⃣ Data Distribution Visualization") selected_column = st.selectbox("Select Column to Visualize", df.columns) fig, ax = plt.subplots() if df[selected_column].dtype in [np.float64, np.int64]: ax.hist(df[selected_column], bins=20) ax.set_title(f"Histogram of {selected_column}") else: df[selected_column].value_counts().plot(kind='bar', ax=ax) ax.set_title(f"Bar Chart of {selected_column}") st.pyplot(fig) # ------------------------------------------------ # Download Processed Data # ------------------------------------------------ st.subheader("⬇ Download Processed Dataset") csv = df.to_csv(index=False).encode('utf-8') st.download_button( label="Download CSV", data=csv, file_name="processed_dataset.csv", mime="text/csv" ) else: st.info("Please upload a CSV file to begin.") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 Data Preprocessing Steps 1️⃣ Handling Missing Values - Drop rows - Fill with Mean / Median / Mode 2️⃣ Encoding Categorical Variables - Label Encoding (assign numbers) - One-Hot Encoding (binary columns) 3️⃣ Scaling Data - Min-Max Normalization → Range [0,1] - Standardization → Mean=0, Std=1 4️⃣ Visualization - Histogram (Numeric Data) - Bar Chart (Categorical Data) --- ### 📊 Why Preprocessing is Important? - Improves model accuracy - Prevents bias due to scale differences - Converts categorical data into machine-readable format """)

📘 Record Explanation (For Submission)

🔹 Missing Value Handling

  • Drop rows OR

  • Fill using mean / median / mode

🔹 Encoding

  • Label Encoding → Suitable for ordinal data

  • One-Hot Encoding → Suitable for nominal data

🔹 Normalization

  • Min-Max Scaling:

X′=X−XminXmax−XminX' = \frac{X - X_{min}}{X_{max} - X_{min}}

🔹 Standardization

Z=X−μσZ = \frac{X - \mu}{\sigma}

🎯 Output Example

Upload dataset →
Select missing strategy →
Apply encoding →
Apply scaling →
Visualize distributions →
Download processed dataset


LAB 5

▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error, r2_score st.set_page_config(page_title="Regression Model", layout="wide") st.title("📈 Lab Program 5: Supervised Learning – Regression Model") st.subheader("Linear Regression with RMSE and R² Evaluation") # ------------------------------------------------ # Upload Dataset # ------------------------------------------------ st.sidebar.header("Upload Dataset") file = st.sidebar.file_uploader("Upload CSV File", type=["csv"]) if file is not None: df = pd.read_csv(file) st.subheader("📌 Dataset Preview") st.write(df.head()) st.write("Shape:", df.shape) # ------------------------------------------------ # Select Target Variable # ------------------------------------------------ st.subheader("🎯 Select Target Variable") target = st.selectbox("Choose Target Column", df.columns) if target: X = df.drop(columns=[target]) y = df[target] # Keep only numeric features X = X.select_dtypes(include=np.number) st.write("Feature Columns Used:", X.columns.tolist()) # ------------------------------------------------ # Train-Test Split # ------------------------------------------------ test_size = st.slider("Select Test Size (%)", 10, 40, 20) / 100 X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=test_size, random_state=42 ) st.success(f"Train Size: {X_train.shape[0]} rows") st.success(f"Test Size: {X_test.shape[0]} rows") # ------------------------------------------------ # Train Linear Regression Model # ------------------------------------------------ if st.button("Train Model"): model = LinearRegression() model.fit(X_train, y_train) y_pred = model.predict(X_test) # ------------------------------------------------ # Evaluation # ------------------------------------------------ rmse = np.sqrt(mean_squared_error(y_test, y_pred)) r2 = r2_score(y_test, y_pred) st.subheader("📊 Model Evaluation") col1, col2 = st.columns(2) with col1: st.metric("RMSE", round(rmse, 4)) with col2: st.metric("R² Score", round(r2, 4)) # ------------------------------------------------ # Visualization # ------------------------------------------------ st.subheader("📉 Actual vs Predicted") fig, ax = plt.subplots() ax.scatter(y_test, y_pred) ax.set_xlabel("Actual Values") ax.set_ylabel("Predicted Values") ax.set_title("Actual vs Predicted") st.pyplot(fig) # ------------------------------------------------ # Model Coefficients # ------------------------------------------------ st.subheader("📌 Model Coefficients") coef_df = pd.DataFrame({ "Feature": X.columns, "Coefficient": model.coef_ }) st.write(coef_df) else: st.info("Please upload a CSV dataset to start.") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 Linear Regression Linear Regression models the relationship between independent variables (X) and dependent variable (Y): \[ Y = β0 + β1X1 + β2X2 + ... + βnXn \] --- ### 🔹 Train-Test Split - Training data → Used to train model - Testing data → Used to evaluate performance - Prevents overfitting --- ### 🔹 Evaluation Metrics 1️⃣ RMSE (Root Mean Squared Error) \[ RMSE = \sqrt{\frac{1}{n} \sum (y_{true} - y_{pred})^2} \] - Lower RMSE → Better model 2️⃣ R² Score (Coefficient of Determination) \[ R^2 = 1 - \frac{SS_{res}}{SS_{tot}} \] - Range: 0 to 1 - Higher R² → Better fit --- ### 📊 Interpretation - R² close to 1 → Good model - Low RMSE → Accurate predictions """)

📘 Sample Viva Questions

  1. What is supervised learning?

  2. What is Linear Regression?

  3. Why do we use Train-Test split?

  4. Difference between RMSE and MSE?

  5. What does R² indicate?


🎯 Output Flow

Upload dataset →
Select target →
Split data →
Train model →
Evaluate using RMSE & R² →
Visualize predictions




LAB 6

🧠 Lab Program 6: Supervised Learning – Classification Model

Logistic Regression & Decision Tree

This app allows you to:

  • ✅ Upload dataset

  • ✅ Select target variable

  • ✅ Choose model (Logistic / Decision Tree)

  • ✅ Train-Test Split

  • ✅ Evaluate using Accuracy, Precision, Recall, F1-score

  • ✅ Visualize Confusion Matrix


▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression from sklearn.tree import DecisionTreeClassifier from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix st.set_page_config(page_title="Classification Model", layout="wide") st.title("🧠 Lab Program 6: Supervised Learning – Classification Model") st.subheader("Logistic Regression & Decision Tree") # ------------------------------------------------ # Upload Dataset # ------------------------------------------------ st.sidebar.header("Upload Dataset") file = st.sidebar.file_uploader("Upload CSV File", type=["csv"]) if file is not None: df = pd.read_csv(file) st.subheader("📌 Dataset Preview") st.write(df.head()) st.write("Shape:", df.shape) # ------------------------------------------------ # Select Target # ------------------------------------------------ st.subheader("🎯 Select Target Variable") target = st.selectbox("Choose Target Column", df.columns) if target: X = df.drop(columns=[target]) y = df[target] # Use only numeric features X = X.select_dtypes(include=np.number) st.write("Feature Columns Used:", X.columns.tolist()) # ------------------------------------------------ # Model Selection # ------------------------------------------------ model_choice = st.selectbox( "Choose Classification Model", ["Logistic Regression", "Decision Tree"] ) # ------------------------------------------------ # Train-Test Split # ------------------------------------------------ test_size = st.slider("Select Test Size (%)", 10, 40, 20) / 100 X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=test_size, random_state=42 ) st.success(f"Train Size: {X_train.shape[0]} rows") st.success(f"Test Size: {X_test.shape[0]} rows") # ------------------------------------------------ # Train Model # ------------------------------------------------ if st.button("Train Model"): if model_choice == "Logistic Regression": model = LogisticRegression(max_iter=1000) else: model = DecisionTreeClassifier() model.fit(X_train, y_train) y_pred = model.predict(X_test) # ------------------------------------------------ # Evaluation Metrics # ------------------------------------------------ accuracy = accuracy_score(y_test, y_pred) precision = precision_score(y_test, y_pred, average='weighted', zero_division=0) recall = recall_score(y_test, y_pred, average='weighted', zero_division=0) f1 = f1_score(y_test, y_pred, average='weighted', zero_division=0) st.subheader("📊 Model Evaluation") col1, col2, col3, col4 = st.columns(4) col1.metric("Accuracy", round(accuracy, 4)) col2.metric("Precision", round(precision, 4)) col3.metric("Recall", round(recall, 4)) col4.metric("F1-Score", round(f1, 4)) # ------------------------------------------------ # Confusion Matrix # ------------------------------------------------ st.subheader("🔎 Confusion Matrix") cm = confusion_matrix(y_test, y_pred) fig, ax = plt.subplots() cax = ax.matshow(cm) plt.title("Confusion Matrix") plt.xlabel("Predicted") plt.ylabel("Actual") plt.colorbar(cax) for (i, j), val in np.ndenumerate(cm): ax.text(j, i, val, ha='center', va='center') st.pyplot(fig) else: st.info("Please upload a CSV dataset to begin.") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 Logistic Regression Used for binary classification problems. Uses sigmoid function: \[ P(Y=1) = \frac{1}{1 + e^{-z}} \] --- ### 🔹 Decision Tree - Tree-based model - Splits data using feature conditions - Easy to interpret --- ### 🔹 Evaluation Metrics 1️⃣ Accuracy \[ Accuracy = \frac{TP + TN}{Total} \] 2️⃣ Precision \[ Precision = \frac{TP}{TP + FP} \] 3️⃣ Recall \[ Recall = \frac{TP}{TP + FN} \] 4️⃣ F1-Score \[ F1 = 2 \times \frac{Precision \times Recall}{Precision + Recall} \] --- ### 🔹 Confusion Matrix | | Predicted Positive | Predicted Negative | |------------|-------------------|-------------------| | Actual Positive | TP | FN | | Actual Negative | FP | TN | """)

📘 Viva Questions

  1. What is classification?

  2. Difference between Logistic Regression and Linear Regression?

  3. When should we use Decision Tree?

  4. What is overfitting?

  5. Why is F1-score important?


🎯 Output Flow

Upload dataset →
Select target →
Choose model →
Split data →
Train model →
Evaluate metrics →
View confusion matrix




LAB 7

Here is a complete Streamlit implementation for:

🔵 Lab Program 7: Unsupervised Learning – Clustering

K-Means Clustering with Elbow Method

This app includes:

  • ✅ Upload real dataset

  • ✅ Select numeric features

  • ✅ Implement K-Means clustering

  • ✅ Elbow Method (Optimal K detection)

  • ✅ Cluster Visualization

  • ✅ Clustered dataset download


▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import pandas as pd import numpy as np import matplotlib.pyplot as plt from sklearn.cluster import KMeans from sklearn.preprocessing import StandardScaler st.set_page_config(page_title="K-Means Clustering", layout="wide") st.title("🔵 Lab Program 7: Unsupervised Learning – Clustering") st.subheader("K-Means Clustering with Elbow Method") # ------------------------------------------------ # Upload Dataset # ------------------------------------------------ st.sidebar.header("Upload Dataset") file = st.sidebar.file_uploader("Upload CSV File", type=["csv"]) if file is not None: df = pd.read_csv(file) st.subheader("📌 Dataset Preview") st.write(df.head()) st.write("Shape:", df.shape) # ------------------------------------------------ # Select Numeric Features # ------------------------------------------------ numeric_cols = df.select_dtypes(include=np.number).columns.tolist() st.subheader("🎯 Select Features for Clustering") selected_features = st.multiselect( "Choose Numeric Features", numeric_cols, default=numeric_cols[:2] if len(numeric_cols) >= 2 else numeric_cols ) if len(selected_features) >= 2: X = df[selected_features] # ------------------------------------------------ # Standardize Data # ------------------------------------------------ scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # ------------------------------------------------ # Elbow Method # ------------------------------------------------ st.subheader("📉 Elbow Method") k_range = range(1, 11) inertia = [] for k in k_range: kmeans = KMeans(n_clusters=k, random_state=42, n_init=10) kmeans.fit(X_scaled) inertia.append(kmeans.inertia_) fig1, ax1 = plt.subplots() ax1.plot(k_range, inertia, marker='o') ax1.set_xlabel("Number of Clusters (K)") ax1.set_ylabel("Inertia (WCSS)") ax1.set_title("Elbow Method") st.pyplot(fig1) # ------------------------------------------------ # Select K # ------------------------------------------------ k_value = st.slider("Select Number of Clusters (K)", 2, 10, 3) if st.button("Run K-Means"): kmeans = KMeans(n_clusters=k_value, random_state=42, n_init=10) clusters = kmeans.fit_predict(X_scaled) df["Cluster"] = clusters st.success("Clustering Completed!") # ------------------------------------------------ # Cluster Visualization (2D) # ------------------------------------------------ st.subheader("📊 Cluster Visualization") fig2, ax2 = plt.subplots() scatter = ax2.scatter( X_scaled[:, 0], X_scaled[:, 1], c=clusters ) ax2.set_xlabel(selected_features[0]) ax2.set_ylabel(selected_features[1]) ax2.set_title("Cluster Plot") st.pyplot(fig2) # ------------------------------------------------ # Show Clustered Data # ------------------------------------------------ st.subheader("📌 Clustered Dataset Preview") st.write(df.head()) # Download Option csv = df.to_csv(index=False).encode('utf-8') st.download_button( label="Download Clustered Dataset", data=csv, file_name="clustered_data.csv", mime="text/csv" ) else: st.warning("Please select at least 2 numeric features.") else: st.info("Please upload a CSV dataset to begin.") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 K-Means Clustering K-Means is an unsupervised learning algorithm that groups data into K clusters. Steps: 1️⃣ Choose number of clusters (K) 2️⃣ Initialize centroids 3️⃣ Assign points to nearest centroid 4️⃣ Update centroids 5️⃣ Repeat until convergence --- ### 🔹 Elbow Method Used to determine optimal K. - Plot K vs Inertia (WCSS) - Look for the "elbow point" - After elbow → diminishing returns --- ### 🔹 Inertia (WCSS) \[ WCSS = \sum (distance\ to\ centroid)^2 \] Lower inertia → tighter clusters --- ### 📊 Applications - Customer Segmentation - Market Analysis - Image Compression - Pattern Recognition """)

📘 Viva Questions

  1. What is unsupervised learning?

  2. Difference between supervised and unsupervised learning?

  3. What is inertia in K-Means?

  4. Why do we use the Elbow method?

  5. What are limitations of K-Means?


🎯 Output Flow

Upload dataset →
Select features →
View elbow graph →
Choose K →
Run K-Means →
Visualize clusters →
Download results




LAB 8

Here is a complete Streamlit implementation for:

🖼️ Lab Program 8: Image Classification using CNN

Convolutional Neural Network (CNN) with TensorFlow/Keras

This app allows you to:

  • ✅ Load image dataset (folder-based structure)

  • ✅ Build CNN architecture

  • ✅ Train model

  • ✅ Evaluate model

  • ✅ Show accuracy & loss curves

  • ✅ Test prediction on new image


📁 Dataset Structure Required

Your dataset folder should look like this:

dataset/ train/ class1/ class2/ test/ class1/ class2/

Example:

dataset/train/cats dataset/train/dogs

▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import tensorflow as tf from tensorflow.keras import layers, models from tensorflow.keras.preprocessing.image import ImageDataGenerator import matplotlib.pyplot as plt import numpy as np from PIL import Image import os st.set_page_config(page_title="CNN Image Classification", layout="wide") st.title("🖼️ Lab Program 8: Image Classification using CNN") st.subheader("Convolutional Neural Network (CNN) Implementation") # ------------------------------------------------ # Dataset Path Input # ------------------------------------------------ st.sidebar.header("Dataset Configuration") train_dir = st.sidebar.text_input("Train Dataset Path", "dataset/train") test_dir = st.sidebar.text_input("Test Dataset Path", "dataset/test") img_size = st.sidebar.slider("Image Size", 32, 128, 64) batch_size = st.sidebar.slider("Batch Size", 8, 64, 16) epochs = st.sidebar.slider("Epochs", 1, 20, 5) # ------------------------------------------------ # Build CNN Model # ------------------------------------------------ def build_cnn_model(input_shape, num_classes): model = models.Sequential([ layers.Conv2D(32, (3,3), activation='relu', input_shape=input_shape), layers.MaxPooling2D((2,2)), layers.Conv2D(64, (3,3), activation='relu'), layers.MaxPooling2D((2,2)), layers.Conv2D(128, (3,3), activation='relu'), layers.MaxPooling2D((2,2)), layers.Flatten(), layers.Dense(128, activation='relu'), layers.Dense(num_classes, activation='softmax') ]) model.compile( optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'] ) return model # ------------------------------------------------ # Train Model # ------------------------------------------------ if st.button("Train CNN Model"): if not os.path.exists(train_dir): st.error("Training directory not found!") else: # Data Generators train_datagen = ImageDataGenerator(rescale=1./255) test_datagen = ImageDataGenerator(rescale=1./255) train_generator = train_datagen.flow_from_directory( train_dir, target_size=(img_size, img_size), batch_size=batch_size, class_mode='categorical' ) test_generator = test_datagen.flow_from_directory( test_dir, target_size=(img_size, img_size), batch_size=batch_size, class_mode='categorical' ) num_classes = len(train_generator.class_indices) model = build_cnn_model((img_size, img_size, 3), num_classes) st.write("### CNN Architecture") model.summary(print_fn=lambda x: st.text(x)) # Training history = model.fit( train_generator, epochs=epochs, validation_data=test_generator ) # Evaluation loss, accuracy = model.evaluate(test_generator) st.success(f"Test Accuracy: {round(accuracy, 4)}") st.success(f"Test Loss: {round(loss, 4)}") # ------------------------------------------------ # Plot Accuracy & Loss # ------------------------------------------------ st.subheader("📊 Training Curves") fig, ax = plt.subplots(1, 2, figsize=(10,4)) ax[0].plot(history.history['accuracy']) ax[0].plot(history.history['val_accuracy']) ax[0].set_title("Accuracy") ax[0].legend(["Train", "Validation"]) ax[1].plot(history.history['loss']) ax[1].plot(history.history['val_loss']) ax[1].set_title("Loss") ax[1].legend(["Train", "Validation"]) st.pyplot(fig) # Save model model.save("cnn_model.h5") st.info("Model saved as cnn_model.h5") # ------------------------------------------------ # Test New Image # ------------------------------------------------ st.subheader("🔍 Test New Image") uploaded_image = st.file_uploader("Upload Image for Prediction", type=["jpg", "png", "jpeg"]) if uploaded_image is not None: image = Image.open(uploaded_image).resize((img_size, img_size)) st.image(image, caption="Uploaded Image") img_array = np.array(image) / 255.0 img_array = np.expand_dims(img_array, axis=0) if os.path.exists("cnn_model.h5"): model = tf.keras.models.load_model("cnn_model.h5") prediction = model.predict(img_array) predicted_class = np.argmax(prediction) st.success(f"Predicted Class Index: {predicted_class}") else: st.warning("Train the model first!") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 Convolutional Neural Network (CNN) CNN is a deep learning model used for image processing. Key Layers: 1️⃣ Convolution Layer - Extracts features - Applies filters 2️⃣ Pooling Layer - Reduces spatial size - Prevents overfitting 3️⃣ Flatten Layer - Converts 2D to 1D 4️⃣ Dense Layer - Fully connected layer - Final classification --- ### 🔹 Loss Function Categorical Cross-Entropy: \[ Loss = -\sum y \log(\hat{y}) \] --- ### 🔹 Evaluation Metrics - Accuracy - Loss --- ### 📊 Applications - Face Recognition - Medical Image Analysis - Object Detection - Autonomous Vehicles """)

📘 Viva Questions

  1. What is CNN?

  2. What is convolution operation?

  3. Why do we use pooling?

  4. What is overfitting in CNN?

  5. Difference between ANN and CNN?


🎯 Output Flow

Set dataset path →
Build CNN →
Train model →
Evaluate accuracy →
View training curves →
Test new image




LAB 9

Here is a complete Streamlit implementation for:

📝 Lab Program 9: NLP – Text Classification / Sentiment Analysis

Sentiment Analysis using Machine Learning (TF-IDF + Logistic Regression)

This app includes:

  • ✅ Text preprocessing (tokenization + stopwords removal)

  • ✅ TF-IDF feature extraction

  • ✅ Logistic Regression model

  • ✅ Train-Test split

  • ✅ Accuracy evaluation

  • ✅ Predict sentiment for new text


▶️ Run the App

Save as app.py and run:

streamlit run app.py

✅ Streamlit Code

import streamlit as st import pandas as pd import numpy as np import re import nltk import matplotlib.pyplot as plt from nltk.corpus import stopwords from nltk.tokenize import word_tokenize from sklearn.model_selection import train_test_split from sklearn.feature_extraction.text import TfidfVectorizer from sklearn.linear_model import LogisticRegression from sklearn.metrics import accuracy_score, classification_report, confusion_matrix # Download NLTK resources (first time only) nltk.download('punkt') nltk.download('stopwords') st.set_page_config(page_title="Sentiment Analysis", layout="wide") st.title("📝 Lab Program 9: NLP – Sentiment Analysis") st.subheader("Text Classification using TF-IDF + Logistic Regression") # ------------------------------------------------ # Text Preprocessing Function # ------------------------------------------------ stop_words = set(stopwords.words("english")) def preprocess_text(text): text = text.lower() text = re.sub(r'[^a-zA-Z\s]', '', text) tokens = word_tokenize(text) tokens = [word for word in tokens if word not in stop_words] return " ".join(tokens) # ------------------------------------------------ # Upload Dataset # ------------------------------------------------ st.sidebar.header("Upload Dataset") file = st.sidebar.file_uploader("Upload CSV (must contain 'text' and 'label' columns)", type=["csv"]) if file is not None: df = pd.read_csv(file) st.subheader("📌 Dataset Preview") st.write(df.head()) if "text" in df.columns and "label" in df.columns: # Preprocess Text st.subheader("🔄 Text Preprocessing") df["clean_text"] = df["text"].apply(preprocess_text) st.write(df[["text", "clean_text"]].head()) # ------------------------------------------------ # TF-IDF Vectorization # ------------------------------------------------ vectorizer = TfidfVectorizer() X = vectorizer.fit_transform(df["clean_text"]) y = df["label"] # ------------------------------------------------ # Train-Test Split # ------------------------------------------------ test_size = st.slider("Select Test Size (%)", 10, 40, 20) / 100 X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=test_size, random_state=42 ) # ------------------------------------------------ # Train Model # ------------------------------------------------ if st.button("Train Model"): model = LogisticRegression(max_iter=1000) model.fit(X_train, y_train) y_pred = model.predict(X_test) # Evaluation accuracy = accuracy_score(y_test, y_pred) st.subheader("📊 Model Evaluation") st.metric("Accuracy", round(accuracy, 4)) # Confusion Matrix st.subheader("🔎 Confusion Matrix") cm = confusion_matrix(y_test, y_pred) fig, ax = plt.subplots() cax = ax.matshow(cm) plt.title("Confusion Matrix") plt.colorbar(cax) for (i, j), val in np.ndenumerate(cm): ax.text(j, i, val, ha='center', va='center') st.pyplot(fig) # Classification Report st.subheader("📋 Classification Report") st.text(classification_report(y_test, y_pred)) # Save model and vectorizer import pickle pickle.dump(model, open("sentiment_model.pkl", "wb")) pickle.dump(vectorizer, open("tfidf_vectorizer.pkl", "wb")) st.success("Model saved successfully!") else: st.error("Dataset must contain 'text' and 'label' columns.") # ------------------------------------------------ # Predict New Text # ------------------------------------------------ st.markdown("---") st.subheader("🔍 Predict Sentiment") user_input = st.text_area("Enter text for sentiment prediction:") if st.button("Predict Sentiment"): import pickle try: model = pickle.load(open("sentiment_model.pkl", "rb")) vectorizer = pickle.load(open("tfidf_vectorizer.pkl", "rb")) clean_input = preprocess_text(user_input) vector_input = vectorizer.transform([clean_input]) prediction = model.predict(vector_input)[0] st.success(f"Predicted Sentiment: {prediction}") except: st.warning("Train the model first!") # ------------------------------------------------ # Theory Section # ------------------------------------------------ st.markdown("---") st.subheader("📘 Theory Explanation") st.markdown(""" ### 🔹 Text Preprocessing Steps 1️⃣ Lowercasing 2️⃣ Remove special characters 3️⃣ Tokenization 4️⃣ Stopword removal --- ### 🔹 TF-IDF (Term Frequency – Inverse Document Frequency) \[ TF-IDF = TF \times IDF \] - Measures importance of words - Reduces impact of common words --- ### 🔹 Logistic Regression for NLP Used for binary/multi-class classification. Applies sigmoid/softmax function to predict class probabilities. --- ### 📊 Applications - Movie review sentiment analysis - Spam detection - Product review classification - Social media analysis """)

📘 Example Dataset Format

Your CSV should look like:

textlabel
I love this productPositive
This is terribleNegative

📘 Viva Questions

  1. What is NLP?

  2. What is tokenization?

  3. What are stopwords?

  4. What is TF-IDF?

  5. Why use Logistic Regression for text classification?


🎯 Output Flow

Upload dataset →
Preprocess text →
Vectorize using TF-IDF →
Train Logistic Regression →
Evaluate accuracy →
Predict new sentiment



LAB 10

Here is a complete End-to-End AI Project (Deployment Included) for:

🚀 Lab Program 10: End-to-End AI Project with Deployment

📈 Option Chosen: Sales Prediction System

This project includes:

  • ✅ Data preprocessing

  • ✅ Model training (Linear Regression)

  • ✅ API deployment using Flask

  • ✅ Monitoring basics

  • ✅ Ethics discussion


📌 Project Overview

We will build:

Sales Prediction System

  • Input: Advertising budget (TV, Radio, Newspaper)

  • Output: Predicted Sales

  • Model: Linear Regression

  • Deployment: Flask REST API


🧠 STEP 1: Data Preprocessing + Model Training (train_model.py)

Use Advertising dataset (CSV format):

TVRadioNewspaperSales

📌 train_model.py

import pandas as pd import numpy as np import pickle from sklearn.model_selection import train_test_split from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error, r2_score # Load dataset df = pd.read_csv("advertising.csv") # Features and Target X = df[["TV", "Radio", "Newspaper"]] y = df["Sales"] # Train-Test Split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # Train Model model = LinearRegression() model.fit(X_train, y_train) # Evaluation y_pred = model.predict(X_test) rmse = np.sqrt(mean_squared_error(y_test, y_pred)) r2 = r2_score(y_test, y_pred) print("RMSE:", rmse) print("R2 Score:", r2) # Save model pickle.dump(model, open("sales_model.pkl", "wb")) print("Model saved successfully!")

Run:

python train_model.py

🌐 STEP 2: Deploy API using Flask

📌 app.py (Flask API)

from flask import Flask, request, jsonify import pickle import numpy as np app = Flask(__name__) # Load trained model model = pickle.load(open("sales_model.pkl", "rb")) @app.route("/") def home(): return "Sales Prediction API Running" @app.route("/predict", methods=["POST"]) def predict(): data = request.get_json(force=True) tv = data["TV"] radio = data["Radio"] newspaper = data["Newspaper"] features = np.array([[tv, radio, newspaper]]) prediction = model.predict(features) return jsonify({ "Predicted Sales": round(float(prediction[0]), 2) }) if __name__ == "__main__": app.run(debug=True)

▶️ Run Flask App

python app.py

API runs at:

http://127.0.0.1:5000/

📌 Test API using Postman / Curl

POST request:

http://127.0.0.1:5000/predict

JSON Body:

{ "TV": 150, "Radio": 20, "Newspaper": 30 }

Response:

{ "Predicted Sales": 18.45 }

📊 STEP 3: Basic Monitoring

You can monitor:

1️⃣ Log Predictions

Add inside predict():

print("Input:", data) print("Prediction:", prediction)

2️⃣ Track Metrics

  • RMSE

  • R²

  • Prediction distribution

  • Error rate over time

3️⃣ Tools for Production Monitoring

  • Prometheus

  • Grafana

  • MLflow

  • Sentry


⚖️ STEP 4: Ethics Discussion

🔹 Bias Risk

  • If dataset is biased → predictions may favor certain markets

🔹 Data Privacy

  • Protect sensitive financial data

🔹 Transparency

  • Explain model decisions

🔹 Over-Reliance Risk

  • AI predictions should support, not replace, human decisions


📘 Architecture Diagram (Conceptual)

Dataset → Preprocessing → Model Training → Save Model ↓ Flask API ↓ Client App

📊 Evaluation Results Example

  • RMSE: 1.65

  • R² Score: 0.89

Interpretation:

  • Model explains 89% variance in sales


📦 Deployment Options

You can deploy:

  • Render

  • Railway

  • Heroku

  • AWS EC2

  • Docker container


📘 Viva Questions

  1. What is model deployment?

  2. Why use Flask for AI APIs?

  3. What is REST API?

  4. How do you monitor ML models?

  5. What are ethical concerns in AI deployment?


🎯 What This Project Demonstrates

StageCompleted
Data Preprocessing✅
Model Training✅
Evaluation✅
API Deployment✅
Monitoring✅
Ethics Discussion✅








Lab11: HR Attrition Calculation

git clone https://github.com/codingacharya/hr-attrition.git

cd hr-attrition

pip install -r requirements.txt

streamlit run app.py


CODE


import streamlit as st
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns

from sklearn.preprocessing import LabelEncoder
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import accuracy_score, confusion_matrix

st.set_page_config(
    page_title="HR Attrition Analytics & Prediction",
    layout="wide"
)

# -----------------------------
# Load Dataset (SAFE)
# -----------------------------
@st.cache_data
def load_data():
    return pd.read_csv("hr_attrition_final_dataset.csv")

df = load_data()

# -----------------------------
# Sidebar Filters
# -----------------------------
st.sidebar.header("🔎 Filters")

dept = st.sidebar.multiselect(
    "Department", df["Department"].unique(), df["Department"].unique()
)
gender = st.sidebar.multiselect(
    "Gender", df["Gender"].unique(), df["Gender"].unique()
)
job = st.sidebar.multiselect(
    "Job Role", df["JobRole"].unique(), df["JobRole"].unique()
)

filtered_df = df[
    (df["Department"].isin(dept)) &
    (df["Gender"].isin(gender)) &
    (df["JobRole"].isin(job))
]

# -----------------------------
# KPI Section
# -----------------------------
st.title("📊 HR Attrition Analytics Dashboard")

total = len(filtered_df)
attrition_count = filtered_df[filtered_df["Attrition"] == "Yes"].shape[0]
attrition_rate = (attrition_count / total) * 100 if total else 0

c1, c2, c3 = st.columns(3)
c1.metric("👥 Total Employees", total)
c2.metric("🚪 Attrition Count", attrition_count)
c3.metric("📉 Attrition Rate (%)", f"{attrition_rate:.2f}")

st.divider()

# -----------------------------
# Visualizations
# -----------------------------
st.subheader("📌 Attrition Distribution")
fig, ax = plt.subplots()
sns.countplot(data=filtered_df, x="Attrition", ax=ax)
st.pyplot(fig)

st.subheader("🏢 Attrition by Department")
fig, ax = plt.subplots()
sns.countplot(data=filtered_df, x="Department", hue="Attrition", ax=ax)
plt.xticks(rotation=30)
st.pyplot(fig)

st.subheader("💰 Income vs Attrition")
fig, ax = plt.subplots()
sns.boxplot(data=filtered_df, x="Attrition", y="MonthlyIncome", ax=ax)
st.pyplot(fig)

st.subheader("😊 Satisfaction Correlation")
satisfaction_cols = [
    "JobSatisfaction",
    "EnvironmentSatisfaction",
    "RelationshipSatisfaction",
    "WorkLifeBalance"
]
fig, ax = plt.subplots()
sns.heatmap(filtered_df[satisfaction_cols].corr(), annot=True, cmap="coolwarm", ax=ax)
st.pyplot(fig)

st.divider()

# -----------------------------
# MACHINE LEARNING
# -----------------------------
st.header("🤖 Attrition Prediction")

ml_df = df.copy()
le = LabelEncoder()

for col in ml_df.select_dtypes(include="object"):
    ml_df[col] = le.fit_transform(ml_df[col])

X = ml_df.drop("Attrition", axis=1)
y = ml_df["Attrition"]

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=42
)

model_type = st.selectbox("Choose Model", ["Logistic Regression", "Random Forest"])

if model_type == "Logistic Regression":
    model = LogisticRegression(max_iter=1000)
else:
    model = RandomForestClassifier(n_estimators=200, random_state=42)

model.fit(X_train, y_train)
y_pred = model.predict(X_test)

acc = accuracy_score(y_test, y_pred)
st.success(f"🎯 Model Accuracy: {acc*100:.2f}%")

st.subheader("📉 Confusion Matrix")
cm = confusion_matrix(y_test, y_pred)
fig, ax = plt.subplots()
sns.heatmap(cm, annot=True, fmt="d", cmap="Blues", ax=ax)
st.pyplot(fig)

if model_type == "Random Forest":
    st.subheader("📌 Feature Importance")
    imp = pd.Series(model.feature_importances_, index=X.columns).sort_values(ascending=False)
    fig, ax = plt.subplots(figsize=(8,5))
    imp.head(10).plot(kind="barh", ax=ax)
    st.pyplot(fig)

st.divider()

# -----------------------------
# Individual Prediction
# -----------------------------
st.header("🧍 Individual Employee Risk Prediction")

input_data = {}
for col in X.columns:
    input_data[col] = st.number_input(col, float(X[col].mean()))

input_df = pd.DataFrame([input_data])
pred = model.predict(input_df)[0]
prob = model.predict_proba(input_df)[0][1]

if pred == 1:
    st.error(f"⚠️ High Attrition Risk ({prob*100:.2f}%)")
else:
    st.success(f"✅ Low Attrition Risk ({prob*100:.2f}%)")

st.divider()

st.subheader("📄 Filtered Data Preview")
st.dataframe(filtered_df, use_container_width=True)





LAB 12: BFS, DFS, Greedy Search and A* search


git clone https://github.com/codingacharya/bfs.git

cd bfs

pip install streamlit networkx matplotlib pandas numpy

streamlit run bfs.py


import streamlit as st
import networkx as nx
import matplotlib.pyplot as plt
import pandas as pd
from collections import deque
import heapq

# ---------------- CONFIG ----------------
st.set_page_config(page_title="Search Algorithms Analytics", layout="wide")

st.title("🔍 Search Algorithms – Visual Analytics Dashboard")
st.markdown("""
This app demonstrates **BFS, DFS, Greedy Best-First, and A\***  
using **graph analytics, heuristics, costs, and visual explanations**.
""")

# ---------------- GRAPH DATA ----------------
edges = [
    ("A", "B", 2),
    ("A", "C", 4),
    ("B", "D", 7),
    ("B", "E", 3),
    ("C", "F", 5),
    ("D", "G", 1),
    ("E", "G", 6),
    ("F", "G", 2),
]

heuristic = {
    "A": 7, "B": 6, "C": 4,
    "D": 1, "E": 3, "F": 2, "G": 0
}

# ---------------- BUILD GRAPH ----------------
G = nx.Graph()
for u, v, w in edges:
    G.add_edge(u, v, weight=w)

pos = nx.spring_layout(G, seed=42)

# ---------------- SEARCH ALGORITHMS ----------------
def bfs(start, goal):
    queue = deque([[start]])
    visited = set()
    steps = []

    while queue:
        path = queue.popleft()
        node = path[-1]
        steps.append(node)

        if node == goal:
            return path, steps

        if node not in visited:
            visited.add(node)
            for neighbor in G.neighbors(node):
                new_path = list(path)
                new_path.append(neighbor)
                queue.append(new_path)

def dfs(start, goal):
    stack = [[start]]
    visited = set()
    steps = []

    while stack:
        path = stack.pop()
        node = path[-1]
        steps.append(node)

        if node == goal:
            return path, steps

        if node not in visited:
            visited.add(node)
            for neighbor in G.neighbors(node):
                new_path = list(path)
                new_path.append(neighbor)
                stack.append(new_path)

def greedy(start, goal):
    pq = [(heuristic[start], [start])]
    visited = set()
    steps = []

    while pq:
        _, path = heapq.heappop(pq)
        node = path[-1]
        steps.append(node)

        if node == goal:
            return path, steps

        if node not in visited:
            visited.add(node)
            for neighbor in G.neighbors(node):
                heapq.heappush(
                    pq, (heuristic[neighbor], path + [neighbor])
                )

def astar(start, goal):
    pq = [(heuristic[start], 0, [start])]
    visited = set()
    steps = []

    while pq:
        f, g, path = heapq.heappop(pq)
        node = path[-1]
        steps.append(node)

        if node == goal:
            return path, steps, g

        if node not in visited:
            visited.add(node)
            for neighbor in G.neighbors(node):
                cost = g + G[node][neighbor]["weight"]
                f_new = cost + heuristic[neighbor]
                heapq.heappush(
                    pq, (f_new, cost, path + [neighbor])
                )

# ---------------- UI CONTROLS ----------------
algo = st.sidebar.selectbox(
    "Choose Algorithm",
    ["BFS", "DFS", "Greedy Best-First", "A*"]
)

start = st.sidebar.selectbox("Start Node", sorted(G.nodes))
goal = st.sidebar.selectbox("Goal Node", sorted(G.nodes), index=6)

# ---------------- RUN ALGORITHM ----------------
if algo == "BFS":
    path, steps = bfs(start, goal)
    total_cost = len(path) - 1

elif algo == "DFS":
    path, steps = dfs(start, goal)
    total_cost = len(path) - 1

elif algo == "Greedy Best-First":
    path, steps = greedy(start, goal)
    total_cost = sum(G[path[i]][path[i+1]]["weight"] for i in range(len(path)-1))

else:
    path, steps, total_cost = astar(start, goal)

# ---------------- VISUALIZATION ----------------
st.subheader("🗺️ Graph Visualization")

plt.figure(figsize=(7, 6))
nx.draw(G, pos, with_labels=True, node_size=2000)
nx.draw_networkx_edge_labels(
    G, pos,
    edge_labels={(u, v): d["weight"] for u, v, d in G.edges(data=True)}
)

path_edges = list(zip(path, path[1:]))
nx.draw_networkx_edges(
    G, pos,
    edgelist=path_edges,
    edge_color="red",
    width=3
)

st.pyplot(plt)

# ---------------- ANALYTICS ----------------
col1, col2 = st.columns(2)

with col1:
    st.subheader("📊 Path Analytics")
    st.metric("Path Found", " → ".join(path))
    st.metric("Total Cost", total_cost)

with col2:
    st.subheader("🧭 Search Behavior")
    st.metric("Nodes Expanded", len(steps))
    st.metric("Unique Nodes", len(set(steps)))

# ---------------- STEP TRACE ----------------
st.subheader("🔎 Step-by-Step Exploration")

df_steps = pd.DataFrame({
    "Step": range(1, len(steps) + 1),
    "Node Expanded": steps,
    "Heuristic h(n)": [heuristic[n] for n in steps]
})

st.dataframe(df_steps, use_container_width=True)

# ---------------- COMPARISON INSIGHTS ----------------
st.subheader("📈 Algorithm Insights")

st.info(f"""
**{algo} Characteristics**

• BFS → Shortest path by steps  
• DFS → Deep exploration  
• Greedy → Fast but heuristic-only  
• A* → Optimal cost-aware search  

**Best for Analytics:**  
• Workflow analysis → BFS  
• Pattern discovery → DFS  
• Quick decisions → Greedy  
• Cost optimization → A*
""")





LAB 13

git clone https://github.com/codingacharya/interactive-sales-dashboard.git

cd interactive-sales-dashboard

pip install streamlit plotly pandas

streamlit run 31.py



import streamlit as st
import pandas as pd
import plotly.express as px
import seaborn as sns
import matplotlib.pyplot as plt

# Load dataset
@st.cache_data
def load_data():
    return pd.read_csv("sales_data.csv", parse_dates=["Date"])

df = load_data()

st.set_page_config(page_title="Sales Dashboard", layout="wide")
st.title("📊 Interactive Sales Dashboard")

# Sidebar Filters
st.sidebar.header("Filters")
regions = st.sidebar.multiselect("Select Region(s)", options=df["Region"].unique(), default=df["Region"].unique())
countries = st.sidebar.multiselect("Select Country(s)", options=df["Country"].unique(), default=df["Country"].unique())
categories = st.sidebar.multiselect("Select Product Category", options=df["ProductCategory"].unique(), default=df["ProductCategory"].unique())
date_range = st.sidebar.date_input("Select Date Range", [df["Date"].min(), df["Date"].max()])

# Apply filters (fix datetime vs date issue)
start_date = pd.to_datetime(date_range[0])
end_date = pd.to_datetime(date_range[1])

df_filtered = df[
    (df["Region"].isin(regions)) &
    (df["Country"].isin(countries)) &
    (df["ProductCategory"].isin(categories)) &
    (df["Date"].between(start_date, end_date))
]

# KPIs
total_sales = df_filtered["TotalSales"].sum()
total_units = df_filtered["UnitsSold"].sum()
avg_price = df_filtered["UnitPrice"].mean()

kpi1, kpi2, kpi3 = st.columns(3)
kpi1.metric("💰 Total Sales", f"${total_sales:,.0f}")
kpi2.metric("📦 Units Sold", f"{total_units:,}")
kpi3.metric("🏷️ Avg Price", f"${avg_price:,.2f}")

st.markdown("---")

# Tabs
tab1, tab2, tab3, tab4, tab5 = st.tabs(["📊 Drill-Down", "📈 Trends", "📊 Aggregations", "📌 Insights", "📄 Data"])

# --- Tab 1: Drill-Down ---
with tab1:
    st.subheader("🔎 Sales Drill-Down")
    level = st.radio("Select Drill Level", ["Region", "Country", "ProductCategory", "Product"], horizontal=True)

    fig = px.bar(
        df_filtered.groupby(level, as_index=False)["TotalSales"].sum(),
        x=level,
        y="TotalSales",
        text="TotalSales",
        title=f"Total Sales by {level}",
        hover_data=["TotalSales"]
    )
    fig.update_traces(texttemplate='%{text:.2s}', textposition='outside')
    st.plotly_chart(fig, use_container_width=True)

# --- Tab 2: Trends ---
with tab2:
    st.subheader("📈 Sales Over Time")
    time_agg = st.selectbox("Aggregate By", ["Day", "Month"], index=1)

    df_time = df_filtered.copy()
    if time_agg == "Month":
        df_time["Period"] = df_time["Date"].dt.to_period("M").dt.to_timestamp()
    else:
        df_time["Period"] = df_time["Date"]

    fig2 = px.line(
        df_time.groupby("Period", as_index=False)["TotalSales"].sum(),
        x="Period",
        y="TotalSales",
        markers=True,
        title=f"Total Sales Over Time ({time_agg})"
    )
    st.plotly_chart(fig2, use_container_width=True)

# --- Tab 3: Aggregations ---
with tab3:
    st.subheader("📊 Aggregation Views")
    agg_dim = st.selectbox("Aggregate Sales By", ["Region", "Country", "ProductCategory", "Product"])

    fig3 = px.pie(
        df_filtered.groupby(agg_dim, as_index=False)["TotalSales"].sum(),
        names=agg_dim,
        values="TotalSales",
        title=f"Sales Distribution by {agg_dim}"
    )
    st.plotly_chart(fig3, use_container_width=True)

# --- Tab 4: Insights (Top N, Correlation, Map) ---
with tab4:
    st.subheader("📌 Insights")

    # Top N Products
    top_n = st.slider("Show Top N Products by Sales", 1, 10, 5)
    top_products = (df_filtered.groupby("Product")["TotalSales"].sum()
                    .sort_values(ascending=False).head(top_n))
    st.bar_chart(top_products)

    # Correlation Heatmap
    st.write("### 🔥 Correlation Heatmap")
    corr = df_filtered[["UnitsSold", "UnitPrice", "TotalSales"]].corr()
    fig_corr, ax = plt.subplots()
    sns.heatmap(corr, annot=True, cmap="Blues", ax=ax)
    st.pyplot(fig_corr)

    # Map Visualization
    st.write("### 🌍 Sales by Country")
    fig_map = px.choropleth(df_filtered,
                            locations="Country",
                            locationmode="country names",
                            color="TotalSales",
                            hover_name="Country",
                            title="Sales by Country")
    st.plotly_chart(fig_map, use_container_width=True)

# --- Tab 5: Data + Export ---
with tab5:
    st.subheader("📄 Filtered Data")
    st.dataframe(df_filtered)

    # Export CSV
    csv = df_filtered.to_csv(index=False).encode("utf-8")
    st.download_button("📥 Download Filtered Data (CSV)", csv, "filtered_sales.csv", "text/csv")



LAB 14



git clone https://github.com/codingacharya/Supply-Chain-Logistics-MCTS.git

cd Supply-Chain-Logistics-MCTS

pip install plotly ortools

streamlit run log.py



"""
Streamlit App: Full-featured Supply Chain / Logistics MCTS Demo

Features included:
- Multiple warehouses
- Multi-truck fleet with capacity constraints
- Customer demands and per-customer price
- Time windows (optional)
- Fuel/energy costs (distance-based + load factor)
- Stochastic traffic (affects travel times/distances during rollouts)
- Order of visiting and delivery allocation decisions via hierarchical MCTS
- Interactive Plotly map visualization

How to run:
1) pip install streamlit numpy plotly
2) streamlit run mcts_supply_chain_full.py

Notes:
- This is a pedagogical demo, not production-grade optimization. For larger instances, integrate a VRP solver (OR-Tools) and replace heavy MCTS with hierarchical approaches.

"""

import streamlit as st
import numpy as np
import random
import math
from dataclasses import dataclass, field
from typing import List, Tuple, Set, Dict, Optional
import plotly.graph_objects as go

# -----------------------------
# Data Classes
# -----------------------------
@dataclass
class Customer:
    id: int
    x: float
    y: float
    demand: int
    price: float
    tw_early: float
    tw_late: float

@dataclass
class Warehouse:
    id: int
    x: float
    y: float
    stock: int

@dataclass
class Truck:
    id: int
    capacity: int
    start_warehouse: int  # warehouse id

@dataclass
class State:
    # assignment partial state: maps truck_id -> ordered list of customers assigned
    assignments: Dict[int, List[int]]
    remaining: Set[int]
    # track load assigned per truck
    loads: Dict[int, int]

    def clone(self):
        return State(assignments={k: list(v) for k, v in self.assignments.items()},
                     remaining=set(self.remaining),
                     loads={k: v for k, v in self.loads.items()})

# -----------------------------
# Utilities
# -----------------------------

def euclid(a: Tuple[float, float], b: Tuple[float, float]) -> float:
    return math.hypot(a[0] - b[0], a[1] - b[1])

# nearest neighbour route given ordered assigned customers (if order empty, we compute greedy order)
def build_route_sequence(start_loc: Tuple[float,float], customers: List[Customer], assigned_ids: List[int]) -> Tuple[List[int], float]:
    id2cust = {c.id: c for c in customers}
    seq = []
    total = 0.0
    current = start_loc
    remaining = assigned_ids[:]
    if not remaining:
        return [], 0.0
    while remaining:
        nxt = min(remaining, key=lambda cid: euclid(current, (id2cust[cid].x, id2cust[cid].y)))
        total += euclid(current, (id2cust[nxt].x, id2cust[nxt].y))
        current = (id2cust[nxt].x, id2cust[nxt].y)
        seq.append(nxt)
        remaining.remove(nxt)
    # return to warehouse
    total += euclid(current, start_loc)
    return seq, total

# compute route distance for a truck starting at warehouse using either assigned order or greedy
def compute_route_distance(warehouse: Warehouse, customers_list: List[Customer], assigned_ids: List[int], traffic_multiplier: float) -> float:
    seq, dist = build_route_sequence((warehouse.x, warehouse.y), customers_list, assigned_ids)
    return dist * traffic_multiplier

# -----------------------------
# MCTS Node and algorithm (action: pick next (truck, customer) pair)
# -----------------------------
class MCTSNode:
    def __init__(self, state: State, parent=None, action: Optional[Tuple[int,int]]=None):
        self.state = state
        self.parent = parent
        self.action = action  # (truck_id, customer_id) or None for root
        self.children: List[MCTSNode] = []
        self.n = 0
        self.w = 0.0
        self.untried: List[Tuple[int,int]] = []  # possible (truck, customer) moves

    def q(self):
        return self.w / self.n if self.n > 0 else 0.0

    def uct_select(self, c=1.4):
        best = None
        best_val = -1e9
        for child in self.children:
            exploitation = child.q()
            exploration = c * math.sqrt(math.log(self.n + 1) / (child.n + 1e-9))
            val = exploitation + exploration
            if val > best_val:
                best_val = val
                best = child
        return best

# rollout: randomly assign remaining customers to trucks then evaluate
def rollout_evaluate(node_state: State, warehouses: Dict[int, Warehouse], trucks: Dict[int, Truck], customers: Dict[int, Customer], cfg) -> float:
    s = node_state.clone()
    # random policy: assign remaining customers to random trucks where capacity allows; if none, assign to nearest warehouse's truck (may exceed capacity -> penalties)
    rem = list(s.remaining)
    random.shuffle(rem)
    truck_ids = list(trucks.keys())
    # naive fill
    for cid in rem:
        # choose trucks that still have capacity
        cand = [tid for tid in truck_ids if s.loads[tid] < trucks[tid].capacity]
        if not cand:
            # force assign to random truck (will cause overload penalty)
            tid = random.choice(truck_ids)
        else:
            tid = random.choice(cand)
        s.assignments[tid].append(cid)
        s.loads[tid] += customers[cid].demand
    # now compute expected reward under stochastic traffic
    traffic_mult = max(0.5, 1.0 + np.random.normal(0, cfg['traffic_sigma']))
    total_reward = evaluate_full_solution(s, warehouses, trucks, customers, cfg, traffic_mult)
    return total_reward

# evaluate full solution: compute revenue minus costs and penalties
def evaluate_full_solution(state: State, warehouses: Dict[int, Warehouse], trucks: Dict[int, Truck], customers: Dict[int, Customer], cfg, traffic_multiplier: float) -> float:
    revenue = 0.0
    travel_cost = 0.0
    time_window_penalty = 0.0
    overload_penalty = 0.0
    unmet_penalty = 0.0

    # For each truck, compute route and deliveries
    for tid, assigned in state.assignments.items():
        truck = trucks[tid]
        wh = warehouses[truck.start_warehouse]
        # build route sequence greedily
        seq, dist = build_route_sequence((wh.x, wh.y), list(customers.values()), assigned)
        # apply traffic multiplier
        route_dist = dist * traffic_multiplier
        # travel cost increases with average load factor
        avg_load = (sum(customers[cid].demand for cid in assigned) / (truck.capacity + 1e-9))
        load_factor = max(0.0, min(1.0, avg_load))
        travel_cost += route_dist * (cfg['travel_cost_per_km'] * (1.0 + cfg['load_cost_factor'] * load_factor))

        # deliver and compute revenue / penalties
        for cid in seq:
            cust = customers[cid]
            qty = min(cust.demand, truck.capacity)  # simplistic: truck may deliver up to capacity but multiple customers reduce capacity
            # For realism, assume delivered equals customer demand if total assigned load <= capacity; else cap
            total_assigned_load = sum(customers[c].demand for c in assigned)
            if total_assigned_load <= truck.capacity:
                delivered = cust.demand
            else:
                # proportionally allocate
                delivered = int(round(cust.demand * (truck.capacity / total_assigned_load)))
            revenue += delivered * cust.price
            if delivered < cust.demand:
                unmet_penalty += (cust.demand - delivered) * cfg['stockout_penalty_per_unit']
            # time windows: compute approximate arrival time: assume speed = cfg['speed_km_per_h'] and time adds from distance; here just approximate using cumulative distance
            # skipping detailed time calc; if time windows enabled, penalize a bit randomly to simulate violation
            if cfg['use_time_windows']:
                # probability of violation increases with traffic
                if random.random() < min(0.3, 0.1 * traffic_multiplier):
                    time_window_penalty += cfg['time_window_violation_penalty']

        # capacity overload
        assigned_load = sum(customers[c].demand for c in assigned)
        if assigned_load > truck.capacity:
            overload_penalty += (assigned_load - truck.capacity) * cfg['overload_penalty_per_unit']

    # warehouses stock constraints are ignored here (assume replenished)

    total = revenue - (travel_cost + time_window_penalty + overload_penalty + unmet_penalty)
    return total

# MCTS planning function that builds assignments
def mcts_plan(warehouses: Dict[int, Warehouse], trucks: Dict[int, Truck], customers: Dict[int, Customer], cfg, iterations:int=400) -> State:
    # initial state: no assignments
    init_assign = {tid: [] for tid in trucks.keys()}
    init_loads = {tid: 0 for tid in trucks.keys()}
    init_state = State(assignments=init_assign, remaining=set(customers.keys()), loads=init_loads)
    root = MCTSNode(init_state)

    # root untried moves: all (truck, customer) possible pairs
    def possible_moves(s: State):
        moves = []
        for tid in trucks.keys():
            for cid in s.remaining:
                moves.append((tid, cid))
        return moves

    root.untried = possible_moves(root.state)

    for it in range(iterations):
        node = root
        # Selection
        while node.untried == [] and node.children:
            node = node.uct_select(c=cfg['uct_c'])
        # Expansion
        if node.untried:
            a = random.choice(node.untried)
            node.untried.remove(a)
            new_state = node.state.clone()
            tid, cid = a
            new_state.assignments[tid].append(cid)
            new_state.loads[tid] += customers[cid].demand
            new_state.remaining.remove(cid)
            child = MCTSNode(new_state, parent=node, action=a)
            child.untried = possible_moves(child.state)
            node.children.append(child)
            node = child
        # Simulation / Rollout
        value = rollout_evaluate(node.state, warehouses, trucks, customers, cfg)
        # Backpropagation
        while node is not None:
            node.n += 1
            node.w += value
            node = node.parent

    # After iterations, pick best child path greedily by visiting children with highest visit count until all customers assigned
    final_state = init_state.clone()
    cur = root
    while final_state.remaining:
        if not cur.children:
            # no explored children; assign remaining randomly
            for cid in list(final_state.remaining):
                tid = random.choice(list(trucks.keys()))
                final_state.assignments[tid].append(cid)
                final_state.loads[tid] += customers[cid].demand
                final_state.remaining.remove(cid)
            break
        # pick child of cur with max visits
        best_child = max(cur.children, key=lambda c: c.n)
        tid, cid = best_child.action
        if cid in final_state.remaining:
            final_state.assignments[tid].append(cid)
            final_state.loads[tid] += customers[cid].demand
            final_state.remaining.remove(cid)
        # move to that child
        cur = best_child
        # if this child has children continue else if still remaining, break to random assign
    return final_state

# -----------------------------
# Streamlit UI & glue
# -----------------------------
st.set_page_config(page_title='MCTS Supply Chain (Full)', layout='wide')
st.title('🚚 Full Supply Chain & Logistics MCTS Demo')

with st.sidebar:
    st.header('Scenario Parameters')
    n_customers = st.slider('Customers', 3, 12, 6)
    n_warehouses = st.slider('Warehouses', 1, 3, 1)
    n_trucks = st.slider('Trucks', 1, 5, 2)
    truck_capacity = st.number_input('Truck capacity (units)', value=30, step=5)
    iterations = st.slider('MCTS iterations', 100, 2000, 600, step=50)

    st.header('Costs & Traffic')
    travel_cost_per_km = st.number_input('Travel cost per km', value=1.0, step=0.1)
    load_cost_factor = st.number_input('Load cost factor (multiplier per load)', value=0.4, step=0.05)
    traffic_sigma = st.slider('Traffic volatility (std dev)', 0.0, 0.8, 0.15, step=0.05)

    st.header('Penalties')
    stockout_penalty = st.number_input('Stockout penalty per unit', value=5.0, step=0.5)
    overload_pen_unit = st.number_input('Overload penalty per unit', value=10.0, step=0.5)
    use_time_windows = st.checkbox('Use time windows', value=True)
    time_window_violation_penalty = st.number_input('Time window violation penalty', value=20.0, step=1.0)

    st.header('Random Seed')
    seed = st.number_input('Seed', value=42, step=1)

# Build scenario
random.seed(int(seed))
np.random.seed(int(seed))

# place warehouses and customers in unit square
warehouses: Dict[int, Warehouse] = {}
for i in range(n_warehouses):
    x, y = np.random.rand(), np.random.rand()
    warehouses[i] = Warehouse(id=i, x=float(x), y=float(y), stock=10000)

customers: Dict[int, Customer] = {}
for i in range(1, n_customers+1):
    x, y = np.random.rand(), np.random.rand()
    demand = int(np.random.poisson(5) + 1)
    price = float(np.random.uniform(8.0, 15.0))
    # time windows centered around random times (in hours) with width
    tw_center = np.random.uniform(8, 17)
    tw_width = np.random.uniform(1.0, 4.0)
    customers[i] = Customer(id=i, x=float(x), y=float(y), demand=demand, price=price, tw_early=tw_center - tw_width/2, tw_late=tw_center + tw_width/2)

# trucks start from random warehouses
trucks: Dict[int, Truck] = {}
for t in range(n_trucks):
    wh_id = random.choice(list(warehouses.keys()))
    trucks[t] = Truck(id=t, capacity=int(truck_capacity), start_warehouse=wh_id)

cfg = {
    'travel_cost_per_km': float(travel_cost_per_km),
    'load_cost_factor': float(load_cost_factor),
    'traffic_sigma': float(traffic_sigma),
    'stockout_penalty_per_unit': float(stockout_penalty),
    'overload_penalty_per_unit': float(overload_pen_unit),
    'use_time_windows': bool(use_time_windows),
    'time_window_violation_penalty': float(time_window_violation_penalty),
    'uct_c': 1.4,
}

st.markdown('### Scenario overview')
col1, col2 = st.columns([2, 1])
with col1:
    # plot nodes
    fig = go.Figure()
    # warehouses
    for wid, wh in warehouses.items():
        fig.add_trace(go.Scatter(x=[wh.x], y=[wh.y], mode='markers+text', marker=dict(size=16, symbol='square', color='red'), text=[f'W{wid}'], textposition='top center'))
    # customers
    for cid, c in customers.items():
        fig.add_trace(go.Scatter(x=[c.x], y=[c.y], mode='markers+text', marker=dict(size=10, color='blue'), text=[f'C{cid} (d={c.demand})'], textposition='top center'))
    fig.update_layout(width=700, height=700, title='Warehouses and Customers (unit square)')
    st.plotly_chart(fig)
with col2:
    st.write('**Warehouses**')
    for wid, wh in warehouses.items():
        st.write(f'W{wid}: ({wh.x:.2f}, {wh.y:.2f}) stock={wh.stock}')
    st.write('**Trucks**')
    for tid, tr in trucks.items():
        st.write(f'T{tid}: cap={tr.capacity}, start=W{tr.start_warehouse}')

# Run MCTS planning
with st.spinner('Running MCTS planning...'):
    plan = mcts_plan(warehouses, trucks, customers, cfg, iterations=int(iterations))

# Evaluate plan deterministically with expected traffic multiplier = 1.0
final_reward = evaluate_full_solution(plan, warehouses, trucks, customers, cfg, traffic_multiplier=1.0)

st.success(f'Planning done — Estimated objective: {final_reward:.2f}')

# Show assignments and simple table
st.subheader('Assignments per Truck')
for tid, assigned in plan.assignments.items():
    st.write(f'Truck {tid} (start W{trucks[tid].start_warehouse}) -> Customers: {assigned} | total load {plan.loads[tid]}')

# Build plot of routes with deterministic traffic
fig2 = go.Figure()
# plot warehouses and customers
for wid, wh in warehouses.items():
    fig2.add_trace(go.Scatter(x=[wh.x], y=[wh.y], mode='markers+text', marker=dict(size=14, symbol='square', color='red'), text=[f'W{wid}'], textposition='bottom center'))
for cid, c in customers.items():
    fig2.add_trace(go.Scatter(x=[c.x], y=[c.y], mode='markers+text', marker=dict(size=10, color='blue'), text=[f'C{cid}\nd={c.demand}'], textposition='top center'))

colors = ['green','orange','purple','brown','magenta','cyan']
for idx, (tid, assigned) in enumerate(plan.assignments.items()):
    truck = trucks[tid]
    wh = warehouses[truck.start_warehouse]
    seq, dist = build_route_sequence((wh.x, wh.y), list(customers.values()), assigned)
    xs = [wh.x]
    ys = [wh.y]
    for cid in seq:
        xs.append(customers[cid].x)
        ys.append(customers[cid].y)
    xs.append(wh.x)
    ys.append(wh.y)
    fig2.add_trace(go.Scatter(x=xs, y=ys, mode='lines+markers', line=dict(color=colors[idx % len(colors)], width=2), name=f'Truck {tid}'))

fig2.update_layout(width=900, height=700, title='Planned Routes (deterministic)')
st.plotly_chart(fig2)

# Allow user to simulate stochastic rollout multiple times and show statistics
st.subheader('Stochastic Simulation (rollouts)')
nsim = st.number_input('Number of rollout simulations', min_value=10, max_value=200, value=50, step=10)

if st.button('Run Rollouts'):
    results = []
    for i in range(int(nsim)):
        tm = max(0.4, 1.0 + np.random.normal(0, cfg['traffic_sigma']))
        val = evaluate_full_solution(plan, warehouses, trucks, customers, cfg, traffic_multiplier=tm)
        results.append(val)
    import statistics
    st.write(f'Rollouts: mean={statistics.mean(results):.2f}, std={statistics.stdev(results):.2f}, min={min(results):.2f}, max={max(results):.2f}')
    st.bar_chart(results)

st.markdown('---')
st.caption('This demo uses MCTS over the (truck,customer) assignment decision space and greedy routing. For larger/real problems, consider OR-Tools for VRP and hierarchical planning.')




LAB 15


save this dataset as sales_data.csv


date,region,product,units_sold,unit_price,cost_per_unit,marketing_spend
2024-01-01,North,Laptop,12,70000,55000,15000
2024-01-02,South,Mobile,30,20000,14000,8000
2024-01-03,East,Tablet,20,30000,22000,10000
2024-01-04,West,Laptop,10,70000,55000,12000
2024-01-05,North,Mobile,25,20000,14000,9000
2024-01-06,South,Tablet,18,30000,22000,7000
2024-01-07,East,Laptop,14,70000,55000,11000
2024-01-08,West,Mobile,28,20000,14000,9500
2024-01-09,North,Tablet,22,30000,22000,8500
2024-01-10,South,Laptop,9,70000,55000,6000



App.py

import streamlit as st
import pandas as pd
import numpy as np
from sklearn.linear_model import LinearRegression
from statsmodels.tsa.arima.model import ARIMA
import time

st.set_page_config(page_title="Sales Intelligence Dashboard", layout="wide")

# -----------------------------
# Real-Time Data Loader
# -----------------------------
@st.cache_data(ttl=10)  # auto-refresh every 10 sec
def load_data():
    df = pd.read_csv("sales_data.csv")
    df["date"] = pd.to_datetime(df["date"])
    df["revenue"] = df["units_sold"] * df["unit_price"]
    df["profit"] = (df["unit_price"] - df["cost_per_unit"]) * df["units_sold"] - df["marketing_spend"]
    return df

df = load_data()

st.title("📊 Advanced Sales Analytics Dashboard")

# -----------------------------
# Interactive Filters
# -----------------------------
st.sidebar.header("🔎 Filters")

region = st.sidebar.multiselect(
    "Select Region", df["region"].unique(), default=df["region"].unique()
)
product = st.sidebar.multiselect(
    "Select Product", df["product"].unique(), default=df["product"].unique()
)

date_range = st.sidebar.date_input(
    "Select Date Range",
    [df["date"].min(), df["date"].max()]
)

filtered_df = df[
    (df["region"].isin(region)) &
    (df["product"].isin(product)) &
    (df["date"].between(pd.to_datetime(date_range[0]), pd.to_datetime(date_range[1])))
]

st.subheader("Filtered Sales Data")
st.dataframe(filtered_df)

# -----------------------------
# Descriptive Analytics
# -----------------------------
st.header("1️⃣ Descriptive Analytics")

col1, col2, col3 = st.columns(3)
col1.metric("Total Revenue (₹)", f"{filtered_df['revenue'].sum():,}")
col2.metric("Total Profit (₹)", f"{filtered_df['profit'].sum():,}")
col3.metric("Units Sold", filtered_df["units_sold"].sum())

# -----------------------------
# Diagnostic Analytics
# -----------------------------
st.header("2️⃣ Diagnostic Analytics")

st.subheader("Profit by Product")
st.bar_chart(filtered_df.groupby("product")["profit"].sum())

st.subheader("Revenue by Region")
st.bar_chart(filtered_df.groupby("region")["revenue"].sum())

# -----------------------------
# Time-Series Forecasting
# -----------------------------
st.header("3️⃣ Predictive Analytics (Time-Series)")

ts_data = (
    filtered_df.groupby("date")["revenue"]
    .sum()
    .asfreq("D")
    .fillna(0)
)

model = ARIMA(ts_data, order=(1, 1, 1))
model_fit = model.fit()

forecast = model_fit.forecast(steps=7)

st.subheader("Next 7 Days Revenue Forecast")
forecast_df = pd.DataFrame({"Forecasted Revenue": forecast})
st.line_chart(forecast_df)

# -----------------------------
# Marketing Spend Prediction
# -----------------------------
st.header("4️⃣ Predictive Analytics (Regression)")

X = df[["marketing_spend"]]
y = df["revenue"]

reg_model = LinearRegression()
reg_model.fit(X, y)

future_spend = st.slider("Future Marketing Spend (₹)", 5000, 25000, 12000)
predicted_revenue = reg_model.predict([[future_spend]])[0]

st.metric("Predicted Revenue (₹)", f"{int(predicted_revenue):,}")

# -----------------------------
# Profit Optimization
# -----------------------------
st.header("5️⃣ Prescriptive Analytics (Profit Optimization)")

spend_range = np.arange(5000, 30000, 1000).reshape(-1, 1)
revenue_preds = reg_model.predict(spend_range)

avg_margin = (df["unit_price"] - df["cost_per_unit"]).mean()
expected_units = revenue_preds / df["unit_price"].mean()

profit_preds = (expected_units * avg_margin) - spend_range.flatten()

optimal_index = np.argmax(profit_preds)

st.success(
    f"💡 Optimal Marketing Spend: ₹{spend_range[optimal_index][0]:,}\n"
    f"💰 Expected Profit: ₹{int(profit_preds[optimal_index]):,}"
)

# -----------------------------
# Footer
# -----------------------------
st.caption("Interview-Ready Project | Time-Series | ML | Optimization | Real-Time Data")

















UNIT 1



Introduction to Artificial Intelligence (AI)

What is Artificial Intelligence?

Artificial Intelligence (AI) is a branch of computer science focused on creating systems that can perform tasks that normally require human intelligence. These tasks include learning, reasoning, problem-solving, perception, understanding natural language, and decision-making.

In simple terms, AI enables machines to think, learn, and act intelligently.


Definitions of Artificial Intelligence

Different researchers have defined AI from various perspectives:

  • John McCarthy (1956):
    “Artificial Intelligence is the science and engineering of making intelligent machines.”

  • Alan Turing:
    Proposed that a machine can be considered intelligent if it can mimic human behavior indistinguishably (Turing Test).

  • Russell & Norvig:
    AI is the study of intelligent agents that perceive their environment and take actions to maximize success.


History of Artificial Intelligence

1. Early Beginnings (1940s–1950s)

  • 1943: McCulloch and Pitts introduced the first artificial neuron model.

  • 1950: Alan Turing proposed the Turing Test.

  • 1956: The term Artificial Intelligence was coined at the Dartmouth Conference.

2. Growth and Optimism (1960s–1970s)

  • Development of early AI programs like problem solvers and game-playing systems.

  • Focus on symbolic AI and rule-based systems.

3. AI Winter (1970s–1990s)

  • Limited computing power and unrealistic expectations led to reduced funding.

  • Progress slowed significantly.

4. Revival and Modern AI (2000s–Present)

  • Rise of machine learning, deep learning, and big data.

  • Breakthroughs in image recognition, speech processing, and autonomous systems.


Applications of Artificial Intelligence

AI is widely used across multiple domains:

1. Healthcare

  • Disease diagnosis and medical imaging

  • Drug discovery

  • Personalized treatment plans

2. Finance

  • Fraud detection

  • Algorithmic trading

  • Credit scoring and risk assessment

3. Education

  • Intelligent tutoring systems

  • Personalized learning platforms

  • Automated grading

4. Transportation

  • Self-driving cars

  • Traffic management systems

  • Route optimization

5. Business & Industry

  • Customer service chatbots

  • Predictive maintenance

  • Supply chain optimization

6. Entertainment & Media

  • Recommendation systems (Netflix, YouTube)

  • Game AI

  • Content generation


Conclusion

Artificial Intelligence has evolved from a theoretical concept to a transformative technology impacting almost every industry. With continuous advancements in computing power and algorithms, AI is shaping the future by improving efficiency, accuracy, and decision-making across domains.





Types of Artificial Intelligence

Artificial Intelligence can be classified into three main types based on capability and intelligence level:


1. Narrow AI (Weak AI)

Definition

Narrow AI is designed to perform a specific task or a narrow set of tasks. It operates under predefined constraints and cannot perform beyond its programmed or trained function.

Characteristics

  • Task-specific

  • No self-awareness

  • Cannot generalize knowledge to other domains

  • Most AI systems today fall under this category

Examples

  • Voice assistants (Siri, Alexa, Google Assistant)

  • Recommendation systems (Netflix, Amazon)

  • Face recognition systems

  • Spam email filters

  • Chatbots

Real-World Status

✅ Currently exists and widely used


2. General AI (Strong AI)

Definition

General AI refers to machines that possess human-level intelligence, capable of understanding, learning, and applying knowledge across different domains, similar to a human being.

Characteristics

  • Can reason, learn, and adapt

  • Transfers knowledge across tasks

  • Capable of autonomous decision-making

  • Exhibits cognitive abilities like humans

Examples

  • Hypothetical human-like robots

  • Machines capable of independent thinking and creativity

Real-World Status

⚠️ Does not yet exist (still under research)


3. Superintelligent AI

Definition

Superintelligent AI surpasses human intelligence in all aspects, including creativity, emotional intelligence, problem-solving, and decision-making.

Characteristics

  • Far exceeds human cognitive abilities

  • Capable of self-improvement

  • May outperform humans in every field

Examples

  • Advanced theoretical AI systems

  • Often discussed in science fiction and future AI ethics research

Real-World Status

❌ Does not exist yet (purely theoretical)


Comparison Table

FeatureNarrow AIGeneral AISuperintelligent AI
ScopeSpecific taskAny intellectual taskAll tasks better than humans
Intelligence LevelLimitedHuman-levelBeyond human
Learning AbilityTask-specificGeneral learningSelf-improving
ExistenceYesNoNo

Summary

  • Narrow AI dominates today’s AI applications

  • General AI represents the next major milestone

  • Superintelligent AI raises future ethical and safety concerns



Intelligent Agents and Environments

What is an Intelligent Agent?

An intelligent agent is an entity that perceives its environment through sensors and acts upon that environment using actuators in order to achieve specific goals.

In simple words:
👉 An intelligent agent senses → thinks → acts.


Definition

According to Russell and Norvig:

An intelligent agent is anything that can be viewed as perceiving its environment through sensors and acting upon that environment through actuators.


Components of an Intelligent Agent

  1. Sensors
    Used to perceive the environment

    • Examples: cameras, microphones, keyboard, temperature sensors

  2. Actuators
    Used to act on the environment

    • Examples: motors, speakers, displays, robotic arms

  3. Agent Program
    Software that maps perceptions to actions

  4. Agent Function
    Mathematical function that determines the best action based on percept history


Agent–Environment Interaction

Environment → Sensors → Agent → Actuators → Environment

The agent continuously interacts with the environment to maximize its performance.


Types of Intelligent Agents

1. Simple Reflex Agent

  • Acts based only on current perception

  • Uses condition–action rules

  • No memory

Example:
Automatic door that opens when it detects a person


2. Model-Based Reflex Agent

  • Maintains an internal state

  • Handles partially observable environments

Example:
Robot vacuum remembering cleaned areas


3. Goal-Based Agent

  • Acts to achieve specific goals

  • Uses planning and decision-making

Example:
Navigation system finding the shortest route


4. Utility-Based Agent

  • Chooses actions that maximize a utility function

  • Considers trade-offs

Example:
Self-driving car balancing safety, speed, and comfort


5. Learning Agent

  • Improves performance over time

  • Learns from experience

Components:

  • Learning element

  • Performance element

  • Critic

  • Problem generator

Example:
Game-playing AI that improves with each match


Environment in AI

An environment is everything external to the agent that it interacts with.


Types of Environments

1. Fully Observable vs Partially Observable

  • Fully observable: Agent has complete information
    Example: Chess

  • Partially observable: Limited information
    Example: Autonomous driving


2. Deterministic vs Stochastic

  • Deterministic: Predictable outcome
    Example: Calculator

  • Stochastic: Uncertain outcomes
    Example: Stock market


3. Episodic vs Sequential

  • Episodic: Each action independent
    Example: Image classification

  • Sequential: Current action affects future states
    Example: Chess


4. Static vs Dynamic

  • Static: Environment doesn’t change
    Example: Crossword puzzle

  • Dynamic: Changes over time
    Example: Traffic system


5. Discrete vs Continuous

  • Discrete: Finite actions/states
    Example: Board games

  • Continuous: Infinite states/actions
    Example: Robot navigation


Performance Measure of an Agent

Defines how successful an agent is in achieving its goal.

Example:

  • Taxi driver agent → safety, speed, comfort, fuel efficiency


Summary

  • Intelligent agents perceive, decide, and act

  • Environments define the complexity of agent behavior

  • Agent design depends on environment type and performance goals




Problem Solving and Search Techniques in AI

Problem Solving in AI

Problem solving in Artificial Intelligence involves finding a sequence of actions that transforms an initial state into a goal state.

A problem is defined by:

  • Initial state

  • Goal state

  • State space

  • Actions (operators)

  • Path cost

Search techniques are used to explore the state space efficiently.


Uninformed Search Techniques

Uninformed (Blind) search algorithms do not use any domain-specific knowledge. They explore the search space without knowing how close a state is to the goal.


1. Breadth-First Search (BFS)

Description

BFS explores nodes level by level, expanding the shallowest nodes first.

Algorithm

  • Uses a queue (FIFO)

  • Start from the root node

  • Expand all neighboring nodes before going deeper

Properties

  • Complete: Yes

  • Optimal: Yes (if all step costs are equal)

  • Time Complexity: O(bᵈ)

  • Space Complexity: O(bᵈ)

(b = branching factor, d = depth of solution)

Advantages

  • Finds shortest path

  • Guaranteed to find solution

Disadvantages

  • High memory usage

  • Slow for large search spaces

Example

  • Finding the shortest path in an unweighted graph


2. Depth-First Search (DFS)

Description

DFS explores nodes by going as deep as possible before backtracking.

Algorithm

  • Uses a stack (LIFO) or recursion

Properties

  • Complete: No (can get stuck in infinite loops)

  • Optimal: No

  • Time Complexity: O(bᵐ)

  • Space Complexity: O(bm)

(m = maximum depth)

Advantages

  • Low memory requirement

  • Simple implementation

Disadvantages

  • May not find shortest path

  • Risk of infinite depth

Example

  • Puzzle solving, maze exploration


Informed Search Techniques

Informed (Heuristic) search algorithms use heuristic functions to guide the search toward the goal efficiently.


3. Greedy Best-First Search

Description

Expands the node that appears closest to the goal, based on heuristic value only.

Evaluation Function

f(n)=h(n)f(n) = h(n)

Where:

  • h(n) = estimated cost from node n to the goal

Properties

  • Complete: No

  • Optimal: No

Advantages

  • Fast and efficient

  • Less memory than BFS

Disadvantages

  • Can be misled by bad heuristics

  • May find suboptimal solutions

Example

  • Route finding using straight-line distance


4. A* Search Algorithm

Description

A* combines actual cost and heuristic cost to find the optimal solution efficiently.

Evaluation Function

f(n)=g(n)+h(n)f(n) = g(n) + h(n)

Where:

  • g(n) = cost from start to node n

  • h(n) = estimated cost from node n to goal

Properties

  • Complete: Yes

  • Optimal: Yes (if heuristic is admissible)

  • Time Complexity: O(bᵈ)

  • Space Complexity: O(bᵈ)

Advantages

  • Finds optimal solution

  • Efficient with good heuristics

Disadvantages

  • High memory usage

  • Performance depends on heuristic quality

Example

  • GPS navigation systems

  • Game AI pathfinding


Comparison Table

AlgorithmTypeCompleteOptimalUses Heuristic
BFSUninformedYesYesNo
DFSUninformedNoNoNo
Greedy Best-FirstInformedNoNoYes
A*InformedYesYesYes

Summary

  • Uninformed search explores blindly

  • Informed search uses heuristics for efficiency

  • A* is the most powerful and widely used search algorithm




Introduction to Knowledge Representation and Reasoning (KRR)

What is Knowledge Representation?

Knowledge Representation (KR) is the field of AI concerned with how knowledge can be formally represented so that a computer system can understand, reason, and make decisions.

It bridges human knowledge and machine reasoning.


Definition

Knowledge Representation is the method used to encode knowledge in a form that an AI system can process to solve complex problems.


Goals of Knowledge Representation

  • Represent real-world information accurately

  • Enable logical reasoning and inference

  • Support decision-making

  • Handle incomplete or uncertain knowledge


Types of Knowledge

  1. Declarative Knowledge – Facts and information
    Example: “Paris is the capital of France”

  2. Procedural Knowledge – How to perform tasks
    Example: Steps to solve a math problem

  3. Heuristic Knowledge – Experience-based rules
    Example: “If traffic is heavy, take an alternate route”

  4. Meta-Knowledge – Knowledge about knowledge
    Example: Knowing which strategy to use


Knowledge Representation Techniques

  1. Logical Representation

    • Propositional Logic

    • First-Order Predicate Logic

  2. Semantic Networks

    • Graph-based representation

    • Nodes represent objects, edges represent relationships

  3. Frames

    • Structured data representation using slots and values

  4. Production Rules

    • IF–THEN rules

    • Used in expert systems


Reasoning in AI

Reasoning is the process of deriving new knowledge from existing facts.

Types of Reasoning

  • Deductive Reasoning: General → Specific

  • Inductive Reasoning: Specific → General

  • Abductive Reasoning: Best explanation inference


Basics of Machine Learning

What is Machine Learning?

Machine Learning (ML) is a subset of AI that enables systems to learn from data and improve performance without being explicitly programmed.


Types of Machine Learning

  1. Supervised Learning

    • Uses labeled data

    • Examples: Linear regression, classification

  2. Unsupervised Learning

    • Uses unlabeled data

    • Examples: Clustering, dimensionality reduction

  3. Reinforcement Learning

    • Learns through reward and punishment

    • Examples: Game-playing agents, robotics


Basic ML Workflow

  1. Data collection

  2. Data preprocessing

  3. Feature extraction

  4. Model training

  5. Evaluation

  6. Prediction


AI vs ML vs DL

Definitions

  • Artificial Intelligence (AI):
    Broad field focused on creating intelligent machines

  • Machine Learning (ML):
    Subset of AI that learns patterns from data

  • Deep Learning (DL):
    Subset of ML using multi-layer neural networks


Comparison Table

FeatureAIMLDL
ScopeBroadestSubset of AISubset of ML
Data DependencyLow to highHighVery high
Feature EngineeringManualMostly manualAutomatic
ComplexityConceptualModerateHigh
ExamplesExpert systems, robotsSpam filtersImage & speech recognition

Relationship Diagram

Artificial Intelligence └── Machine Learning └── Deep Learning

Key Differences at a Glance

  • AI focuses on decision-making

  • ML focuses on learning from data

  • DL focuses on learning from large data using neural networks


Summary

  • Knowledge Representation enables machines to store and reason with knowledge

  • Reasoning helps AI systems draw conclusions

  • ML allows systems to learn from experience

  • DL powers modern AI breakthroughs



UNIT 2


Introduction to Machine Learning

What is Machine Learning?

Machine Learning (ML) is a branch of Artificial Intelligence that enables computer systems to learn patterns from data and make predictions or decisions without being explicitly programmed.

ML systems improve their performance as they are exposed to more data.


Types of Machine Learning

Machine Learning is broadly classified into:

  1. Supervised Learning

  2. Unsupervised Learning

  3. Reinforcement Learning (advanced topic)

This section focuses on supervised and unsupervised learning.


Supervised Learning

Definition

Supervised learning is a type of ML where the model is trained using labeled data, meaning each input has a corresponding known output.


Main Tasks in Supervised Learning

  • Regression

  • Classification


1. Regression

Definition

Regression is used when the output variable is continuous.

Goal

To predict a numerical value based on input features.

Common Algorithms

  • Linear Regression

  • Polynomial Regression

  • Ridge and Lasso Regression

Examples

  • House price prediction

  • Temperature forecasting

  • Salary prediction


2. Classification

Definition

Classification is used when the output variable is categorical or discrete.

Goal

To assign an input to one of the predefined classes.

Common Algorithms

  • Logistic Regression

  • Decision Trees

  • k-Nearest Neighbors (KNN)

  • Support Vector Machines (SVM)

  • Naive Bayes

Examples

  • Spam vs non-spam email detection

  • Disease diagnosis (positive/negative)

  • Image classification


Unsupervised Learning

Definition

Unsupervised learning works with unlabeled data. The system tries to discover hidden patterns or structures in the data.


Main Tasks in Unsupervised Learning

  • Clustering

  • Dimensionality Reduction


1. Clustering

Definition

Clustering groups similar data points into clusters based on their features.

Goal

To identify natural groupings in data.

Common Algorithms

  • K-Means

  • Hierarchical Clustering

  • DBSCAN

Examples

  • Customer segmentation

  • Document clustering

  • Market analysis


2. Dimensionality Reduction

Definition

Dimensionality reduction reduces the number of input features while preserving important information.

Goal

  • Reduce computational complexity

  • Remove noise

  • Improve visualization

Common Techniques

  • Principal Component Analysis (PCA)

  • Linear Discriminant Analysis (LDA)

  • t-SNE

Examples

  • Data visualization

  • Feature selection

  • Image compression


Supervised vs Unsupervised Learning

FeatureSupervised LearningUnsupervised Learning
DataLabeledUnlabeled
OutputKnownUnknown
GoalPredictionPattern discovery
ExamplesRegression, ClassificationClustering, Dimensionality Reduction

Summary

  • Machine Learning enables systems to learn from data

  • Supervised learning predicts known outputs

  • Regression handles numerical values

  • Classification handles categories

  • Unsupervised learning finds hidden structures

  • Clustering groups data

  • Dimensionality reduction simplifies data



Feature Engineering and Preprocessing

What is Feature Engineering?

Feature engineering is the process of selecting, creating, transforming, and optimizing input features to improve the performance of a machine learning model.

👉 Better features often matter more than better algorithms.


What is Data Preprocessing?

Data preprocessing is the step where raw data is cleaned, formatted, and prepared before feeding it into a machine learning model.

Feature engineering and preprocessing are closely related and usually performed together.


Importance of Feature Engineering & Preprocessing

  • Improves model accuracy

  • Reduces overfitting

  • Speeds up training

  • Handles noisy and incomplete data

  • Makes patterns more learnable


Data Preprocessing Steps

1. Data Cleaning

Handling errors and inconsistencies in data.

a) Handling Missing Values

  • Remove rows or columns

  • Fill with mean, median, or mode

  • Use model-based imputation

Example:
Missing age → replace with average age


b) Handling Outliers

  • Detect using box plots or Z-score

  • Remove or cap extreme values

Example:
Extremely high salary values skewing results


2. Data Transformation

a) Encoding Categorical Data

Convert non-numeric data into numeric form.

  • Label Encoding – Assign numbers to categories

  • One-Hot Encoding – Create binary columns

Example:
Gender → Male = 0, Female = 1


b) Feature Scaling

Ensures features are on the same scale.

  • Normalization (Min-Max Scaling)

  • Standardization (Z-score scaling)

Why needed?
Algorithms like KNN, SVM, and Gradient Descent are scale-sensitive.


3. Data Reduction

Reducing size without losing important information.

  • Feature selection

  • Dimensionality reduction (PCA)


Feature Engineering Techniques

1. Feature Creation

Creating new features from existing ones.

Example:

  • Date → Day, Month, Year

  • Height & Weight → BMI


2. Feature Selection

Selecting the most relevant features.

Methods

  • Filter methods (correlation, chi-square)

  • Wrapper methods (forward selection)

  • Embedded methods (Lasso regression)


3. Feature Transformation

Changing feature distribution to improve learning.

  • Log transformation

  • Square root transformation

  • Power transformation


4. Handling Imbalanced Data

  • Oversampling (SMOTE)

  • Undersampling

  • Class weight adjustment


Preprocessing Pipeline

Typical ML preprocessing flow:

Raw Data ↓ Data Cleaning ↓ Encoding & Scaling ↓ Feature Engineering ↓ Model Training

Examples

  • Spam detection: Text vectorization (TF-IDF)

  • House price prediction: Location encoding, area normalization

  • Image data: Resizing, normalization


Summary

  • Feature engineering improves what the model learns

  • Preprocessing improves how the model learns

  • Good features + clean data = better models

  • Essential step before any ML algorithm



Evaluation Metrics in Machine Learning

Evaluation metrics are used to measure how well a machine learning model performs on unseen data.

They differ based on the type of problem:

  • Classification → Accuracy, Precision, Recall, F1-Score

  • Regression → RMSE


Confusion Matrix (for Classification)

Actual \ PredictedPositiveNegative
PositiveTP (True Positive)FN (False Negative)
NegativeFP (False Positive)TN (True Negative)

This matrix forms the basis for most classification metrics.


1. Accuracy

Definition

Accuracy measures the overall correctness of the model.

Formula

Accuracy=TP+TNTP+TN+FP+FN\text{Accuracy} = \frac{TP + TN}{TP + TN + FP + FN}

When to Use

  • Balanced datasets

  • Equal cost of FP and FN

Limitation

  • Misleading for imbalanced datasets


2. Precision

Definition

Precision measures how many predicted positives are actually correct.

Formula

Precision=TPTP+FP\text{Precision} = \frac{TP}{TP + FP}

When to Use

  • When false positives are costly

Example

  • Spam detection (don’t mark important emails as spam)


3. Recall (Sensitivity)

Definition

Recall measures how many actual positives are correctly identified.

Formula

Recall=TPTP+FN\text{Recall} = \frac{TP}{TP + FN}

When to Use

  • When false negatives are costly

Example

  • Disease detection (don’t miss sick patients)


4. F1-Score

Definition

F1-score is the harmonic mean of Precision and Recall, balancing both.

Formula

F1=2×Precision×RecallPrecision+Recall\text{F1} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}

When to Use

  • Imbalanced datasets

  • Need balance between precision and recall


Regression Evaluation Metric

5. RMSE (Root Mean Square Error)

Definition

RMSE measures the average magnitude of prediction errors in regression models.

Formula

RMSE=1n∑i=1n(yi−y^i)2\text{RMSE} = \sqrt{\frac{1}{n} \sum_{i=1}^{n} (y_i - \hat{y}_i)^2}

Where:

  • yiy_i = actual value

  • y^i\hat{y}_i = predicted value

Characteristics

  • Penalizes large errors more

  • Same unit as target variable

When to Use

  • Continuous value prediction

  • When large errors matter


Metric Comparison Summary

MetricProblem TypeBest Used When
AccuracyClassificationBalanced data
PrecisionClassificationFP is costly
RecallClassificationFN is costly
F1-ScoreClassificationImbalanced data
RMSERegressionLarge errors matter

Quick Tip for Exams

  • Accuracy → overall performance

  • Precision → correctness of positive predictions

  • Recall → completeness of positive detection

  • F1 → balance between precision & recall

  • RMSE → average prediction error




Data Analytics Concepts

What is Data Analytics?

Data Analytics is the process of collecting, cleaning, analyzing, and visualizing data to discover patterns, gain insights, and support decision-making.


1. Data Collection

Definition

Data collection is the process of gathering raw data from various sources for analysis.

Sources of Data

  • Primary sources: Surveys, interviews, sensors, experiments

  • Secondary sources: Databases, websites, reports, APIs, logs

Types of Data

  • Structured: Tables, databases

  • Unstructured: Text, images, videos

  • Semi-structured: JSON, XML

Importance

  • Quality analysis depends on quality data

  • Ensures relevance and accuracy


2. Data Cleaning

Definition

Data cleaning (data cleansing) is the process of identifying and correcting errors, inconsistencies, and missing values in data.

Common Data Issues

  • Missing values

  • Duplicate records

  • Inconsistent formats

  • Outliers and noise

Data Cleaning Techniques

  • Removing or imputing missing values

  • Removing duplicates

  • Standardizing formats

  • Handling outliers

Benefits

  • Improves data quality

  • Increases model accuracy

  • Reduces bias and errors


3. Data Visualization

Definition

Data visualization is the graphical representation of data to communicate insights clearly and effectively.

Common Visualization Tools

  • Bar charts

  • Line charts

  • Pie charts

  • Histograms

  • Heatmaps

Visualization Tools & Libraries

  • Excel, Tableau, Power BI

  • Python (Matplotlib, Seaborn)

Importance

  • Simplifies complex data

  • Identifies trends and patterns

  • Aids decision-making


Types of Data Analytics

Data analytics is classified into four major types based on the question they answer.


1. Descriptive Analytics

Question Answered

👉 What happened?

Description

Summarizes historical data to understand past performance.

Techniques

  • Mean, median, mode

  • Percentages

  • Data aggregation

  • Reports and dashboards

Example

  • Monthly sales reports

  • Website traffic statistics


2. Diagnostic Analytics

Question Answered

👉 Why did it happen?

Description

Analyzes data to identify causes and reasons behind events.

Techniques

  • Drill-down analysis

  • Data correlation

  • Root cause analysis

Example

  • Why sales dropped last quarter

  • Reasons for customer churn


3. Predictive Analytics

Question Answered

👉 What is likely to happen?

Description

Uses historical data and statistical/ML models to predict future outcomes.

Techniques

  • Regression

  • Classification

  • Time-series analysis

  • Machine learning

Example

  • Sales forecasting

  • Demand prediction

  • Credit risk analysis


4. Prescriptive Analytics

Question Answered

👉 What should we do?

Description

Recommends actions based on predictions and constraints.

Techniques

  • Optimization models

  • Simulation

  • Decision rules

  • Reinforcement learning

Example

  • Dynamic pricing strategies

  • Inventory optimization

  • Route planning


Comparison Table

TypeQuestionFocus
DescriptiveWhat happened?Past
DiagnosticWhy did it happen?Causes
PredictiveWhat will happen?Future
PrescriptiveWhat should be done?Action

Summary

  • Data analytics converts raw data into insights

  • Data collection, cleaning, and visualization are foundational steps

  • The four analytics types progress from insight → understanding → prediction → action




Tools: Python Libraries for Data Analytics & ML

Python is the most popular language for data analysis and machine learning because of its powerful, easy-to-use libraries.


1. NumPy (Numerical Python)

Purpose

NumPy is used for numerical computing and mathematical operations in Python.

Key Features

  • N-dimensional arrays (ndarray)

  • Fast mathematical and statistical functions

  • Linear algebra, random numbers

Common Uses

  • Matrix operations

  • Scientific computing

  • Base library for Pandas, Scikit-learn

Example

import numpy as np arr = np.array([1, 2, 3]) print(arr.mean())

2. Pandas

Purpose

Pandas is used for data manipulation and analysis.

Key Features

  • Data structures: Series, DataFrame

  • Handling missing data

  • Filtering, grouping, merging datasets

Common Uses

  • Data cleaning

  • Exploratory Data Analysis (EDA)

  • CSV, Excel, SQL data handling

Example

import pandas as pd df = pd.read_csv("data.csv") print(df.head())

3. Matplotlib

Purpose

Matplotlib is used for basic data visualization.

Key Features

  • Highly customizable plots

  • Supports line, bar, scatter, histogram plots

Common Uses

  • Trend analysis

  • Visualizing numerical data

Example

import matplotlib.pyplot as plt plt.plot([1, 2, 3], [4, 5, 6]) plt.show()

4. Seaborn

Purpose

Seaborn is built on Matplotlib and provides statistical and attractive visualizations.

Key Features

  • High-level plotting functions

  • Built-in themes

  • Works well with Pandas DataFrames

Common Uses

  • Heatmaps

  • Distribution plots

  • Relationship analysis

Example

import seaborn as sns sns.histplot([1, 2, 2, 3, 4])

5. Scikit-learn (sklearn)

Purpose

Scikit-learn is used for machine learning model building and evaluation.

Key Features

  • Classification, regression, clustering algorithms

  • Feature selection and preprocessing

  • Model evaluation metrics

Common Uses

  • Train ML models

  • Model validation

  • Pipelines and cross-validation

Example

from sklearn.linear_model import LinearRegression model = LinearRegression() model.fit(X, y)

Comparison Table

LibraryMain Use
NumPyNumerical computing
PandasData manipulation
MatplotlibBasic visualization
SeabornStatistical visualization
Scikit-learnMachine learning

Typical Data Analytics Workflow Using These Tools

Data Collection → Pandas Data Cleaning → Pandas + NumPy Visualization → Matplotlib + Seaborn Model Building → Scikit-learn Evaluation → Scikit-learn

Summary

  • NumPy handles numbers efficiently

  • Pandas manages structured data

  • Matplotlib & Seaborn visualize insights

  • Scikit-learn builds and evaluates ML models

These libraries together form the core toolkit of a data analyst or ML engineer.



UNIT 3


🌱 What is Deep Learning?

Deep Learning is a subset of Machine Learning that uses neural networks with many layers (hence deep) to learn patterns from data automatically.

It shines at tasks like:

  • Image & face recognition

  • Speech recognition

  • Language translation

  • Medical diagnosis

  • Autonomous driving


🧠 Neural Networks (Big Picture)

A Neural Network is inspired by the human brain.

It consists of:

  1. Input Layer – receives data

  2. Hidden Layers – extract patterns

  3. Output Layer – gives prediction

Each layer contains neurons, and neurons are connected by weights.

Input → Hidden Layer(s) → Output

⚙️ Perceptron (The Simplest Neuron)

The Perceptron is the basic building block of a neural network.

How it works:

  1. Multiply inputs by weights

  2. Add bias

  3. Apply activation function

Mathematical form:

y=f(∑wixi+b)y = f(\sum w_i x_i + b)

Where:

  • xx = input

  • ww = weight

  • bb = bias

  • ff = activation function

👉 A single perceptron can solve linearly separable problems (like AND, OR), but not XOR.


🔥 Activation Functions

Activation functions introduce non-linearity, allowing neural networks to learn complex patterns.

Common Activation Functions:

1️⃣ Sigmoid

f(x)=11+e−xf(x) = \frac{1}{1 + e^{-x}}
  • Output range: (0,1)

  • Used in binary classification

  • ❌ Vanishing gradient problem


2️⃣ ReLU (Rectified Linear Unit)

f(x)=max⁡(0,x)f(x) = \max(0, x)
  • Fast & efficient

  • Most popular in hidden layers

  • ❌ Dying ReLU problem


3️⃣ Tanh

f(x)=tanh⁡(x)f(x) = \tanh(x)
  • Output range: (-1,1)

  • Better than sigmoid in many cases


4️⃣ Softmax

f(xi)=exi∑exjf(x_i) = \frac{e^{x_i}}{\sum e^{x_j}}
  • Used in multi-class classification

  • Output is probability distribution


🔄 Forward Propagation

Forward propagation is how data moves from input to output.

Steps:

  1. Input data enters the network

  2. Weighted sum + bias calculated

  3. Activation function applied

  4. Output is generated

Example:

Input → Weighted Sum → Activation → Output

This output is compared with the actual value to calculate loss.


📉 Loss Function

Measures how wrong the prediction is.

Examples:

  • Mean Squared Error (Regression)

  • Binary Cross-Entropy (Binary classification)

  • Categorical Cross-Entropy (Multi-class)


🔁 Backpropagation

Backpropagation is how the network learns.

Goal:

Minimize loss by updating weights.

Steps:

  1. Compute loss

  2. Calculate gradients using chain rule

  3. Update weights using Gradient Descent

w=w−η∂L∂ww = w - \eta \frac{\partial L}{\partial w}

Where:

  • η\eta = learning rate

  • LL = loss function

👉 Backpropagation moves from output layer to input layer.


⚡ Gradient Descent

Optimizes the weights by moving in the direction of minimum loss.

Types:

  • Batch Gradient Descent

  • Stochastic Gradient Descent (SGD)

  • Mini-batch Gradient Descent


🧩 Why Deep Networks Work

  • Multiple layers learn hierarchical features

  • Early layers → simple patterns

  • Deeper layers → complex representations

Example (Image):

Edges → Shapes → Objects → Meaning

✅ Summary

  • Perceptron: basic neuron

  • Neural Network: stack of perceptrons

  • Activation functions: add non-linearity

  • Forward propagation: prediction

  • Backpropagation: learning from errors



🖼️ What is Image Analytics?

Image analytics is the process of extracting meaningful information from images using algorithms and models.

Common tasks:

  • Image classification

  • Object detection

  • Face recognition

  • Medical image analysis

  • Satellite image interpretation

👉 CNNs are designed specifically for image data.


🧠 Why Not Normal Neural Networks?

Images have:

  • High dimensionality (e.g., 224×224×3)

  • Spatial relationships between pixels

Fully connected networks:

  • Have too many parameters ❌

  • Ignore spatial structure ❌

CNNs solve this using local connectivity and weight sharing ✅


🧱 What is a CNN?

A Convolutional Neural Network is a deep learning model that automatically learns spatial features from images using convolution operations.

Typical CNN Architecture:

Input Image ↓ Convolution Layer ↓ Activation (ReLU) ↓ Pooling Layer ↓ Fully Connected Layer ↓ Output

🔍 Convolution Operation

Convolution uses a filter (kernel) to scan the image.

How it works:

  • A small matrix (e.g., 3×3) slides over the image

  • Element-wise multiplication + sum

  • Produces a feature map

👉 Filters learn to detect:

  • Edges

  • Corners

  • Textures

  • Shapes

Mathematical form:

Feature=(Image∗Kernel)+BiasFeature = (Image * Kernel) + Bias

🎛️ Important CNN Components

1️⃣ Convolution Layer

  • Extracts features

  • Parameters:

    • Kernel size (3×3, 5×5)

    • Stride

    • Padding

    • Number of filters

Output size:

(W−F+2P)S+1\frac{(W - F + 2P)}{S} + 1

2️⃣ Activation Function (ReLU)

f(x)=max⁡(0,x)f(x) = \max(0, x)
  • Introduces non-linearity

  • Speeds up training


3️⃣ Pooling Layer

Reduces spatial dimensions while keeping important features.

Types:

  • Max Pooling (most common)

  • Average Pooling

Benefits:

  • Reduces computation

  • Prevents overfitting

  • Adds translation invariance


4️⃣ Fully Connected Layer

  • Flattens feature maps

  • Performs final classification

  • Works like a traditional neural network


5️⃣ Softmax Output Layer

Used for multi-class image classification:

P(yi)=ezi∑ezjP(y_i) = \frac{e^{z_i}}{\sum e^{z_j}}

🔁 Forward & Backpropagation in CNN

Forward Propagation:

  1. Image → Convolution

  2. ReLU

  3. Pooling

  4. FC layers

  5. Prediction

Backpropagation:

  • Loss gradient flows backward

  • Updates:

    • Filter weights

    • Biases

  • Uses Gradient Descent / Adam

👉 CNNs learn filters automatically.


📉 Loss Functions in CNNs

  • Categorical Cross-Entropy → multi-class

  • Binary Cross-Entropy → binary classification

  • IoU / Dice Loss → segmentation


🧬 Popular CNN Architectures

  • LeNet-5 – handwritten digits

  • AlexNet – ImageNet breakthrough

  • VGG-16 / VGG-19 – deep but simple

  • ResNet – skip connections

  • Inception – multi-scale filters


🧠 Why CNNs Work So Well

  • Local feature extraction

  • Parameter sharing

  • Hierarchical learning

  • Robust to translation & noise

Example:

Pixels → Edges → Shapes → Objects → Class

🏥 Applications of CNNs in Image Analytics

  • Medical imaging (tumor detection)

  • Facial recognition systems

  • Autonomous vehicles

  • Remote sensing & satellite imagery

  • Quality inspection in manufacturing


✅ Summary

  • CNNs are specialized for image data

  • Convolution extracts features

  • Pooling reduces size

  • Fully connected layers classify

  • Backpropagation trains filters




⏳ What is Sequential Data?

Sequential data has an order, and past values matter.

Examples:

  • Text & sentences

  • Speech & audio signals

  • Time-series (stock prices, weather, sensor data)

  • DNA sequences

👉 Traditional neural networks ignore order — RNNs don’t.


🔁 Recurrent Neural Networks (RNN)

An RNN is a neural network with loops, allowing information to persist across time steps.

Key idea:

The output at time t depends on:

  • Current input xₜ

  • Previous hidden state hₜ₋₁

x₁ → h₁ → y₁ ↓ x₂ → h₂ → y₂ ↓ x₃ → h₃ → y₃

🧮 RNN Mathematical Form

ht=f(Whht−1+Wxxt+b)h_t = f(W_h h_{t-1} + W_x x_t + b)
yt=g(Wyht)y_t = g(W_y h_t)

Where:

  • hₜ = hidden state (memory)

  • f = activation (tanh / ReLU)

  • g = output activation (softmax, sigmoid)


🔄 Types of RNN Architectures

  • One-to-One → simple NN

  • One-to-Many → image captioning

  • Many-to-One → sentiment analysis

  • Many-to-Many → machine translation


⚠️ Problems with Basic RNNs

1️⃣ Vanishing Gradient

  • Gradients shrink during backpropagation

  • Model forgets long-term dependencies

2️⃣ Exploding Gradient

  • Gradients grow too large

  • Training becomes unstable

👉 This is why LSTM was introduced.


🧠 Long Short-Term Memory (LSTM)

LSTM is a special type of RNN designed to remember long-term information.

Core idea:

  • Uses gates to control information flow

  • Decides what to remember, forget, and output


🚪 LSTM Gates Explained

1️⃣ Forget Gate

Decides what to remove from memory.

ft=σ(Wf[ht−1,xt]+bf)f_t = \sigma(W_f [h_{t-1}, x_t] + b_f)

2️⃣ Input Gate

Decides what new info to store.

it=σ(Wi[ht−1,xt]+bi)i_t = \sigma(W_i [h_{t-1}, x_t] + b_i)
C~t=tanh⁡(Wc[ht−1,xt]+bc)\tilde{C}_t = \tanh(W_c [h_{t-1}, x_t] + b_c)

3️⃣ Cell State Update

Ct=ft⋅Ct−1+it⋅C~tC_t = f_t \cdot C_{t-1} + i_t \cdot \tilde{C}_t

4️⃣ Output Gate

Decides what to output.

ot=σ(Wo[ht−1,xt]+bo)o_t = \sigma(W_o [h_{t-1}, x_t] + b_o)
ht=ot⋅tanh⁡(Ct)h_t = o_t \cdot \tanh(C_t)

👉 Cₜ acts as long-term memory


🔁 Backpropagation Through Time (BPTT)

RNNs & LSTMs are trained using BPTT:

  • Network is unrolled over time

  • Errors flow backward across time steps

  • Weights updated using gradient descent


📉 Loss Functions

  • Cross-Entropy → text, classification

  • MSE / MAE → time-series prediction


🆚 RNN vs LSTM

FeatureRNNLSTM
Long-term memory❌ Poor✅ Strong
Vanishing gradient❌ Yes✅ Reduced
ComplexityLowHigh
PerformanceModerateExcellent

🚀 Applications of RNN & LSTM

  • Language translation

  • Speech recognition

  • Text generation

  • Stock price forecasting

  • Anomaly detection in sensor data


🧠 Intuition in One Line

  • RNN: short memory

  • LSTM: smart memory with gates


✅ Summary

  • RNNs process data sequentially

  • Hidden state carries information

  • LSTMs solve long-term dependency problems

  • Gates control memory flow




🗣️ What is Natural Language Processing (NLP)?

Natural Language Processing (NLP) is a field of AI that enables machines to understand, interpret, and generate human language.

It combines:

  • Computer Science

  • Linguistics

  • Machine Learning / Deep Learning


🧠 Why NLP is Hard

Human language is:

  • Ambiguous (“bank”, “bat”)

  • Context-dependent

  • Unstructured

  • Full of slang, sarcasm, and errors

👉 NLP converts text into a structured, numerical form machines can work with.


🧩 NLP Pipeline (Core Steps)

Raw Text ↓ Text Cleaning ↓ Tokenization ↓ Text Representation ↓ Modeling ↓ Evaluation

🧹 Text Preprocessing

Cleaning text improves model performance.

Common Steps:

  • Lowercasing

  • Removing punctuation & special characters

  • Removing stopwords (is, the, and)

  • Handling contractions (don't → do not)

  • Stemming / Lemmatization


✂️ Tokenization

Tokenization splits text into smaller units.

Types:

  • Word Tokenization

    • “I love NLP” → ["I", "love", "NLP"]

  • Sentence Tokenization

  • Subword Tokenization (BPE, WordPiece)

  • Character Tokenization


🌱 Stemming vs Lemmatization

FeatureStemmingLemmatization
ApproachRule-basedDictionary-based
OutputRoot wordMeaningful word
Example“running → run”“better → good”

🔢 Text Representation (Vectorization)

Machines understand numbers, not words.

1️⃣ Bag of Words (BoW)

  • Counts word frequency

  • Ignores word order

Example:

"I love NLP" → [1, 1, 1]

2️⃣ TF-IDF

Balances frequency with importance.

TF-IDF=TF×log⁡NDFTF\text{-}IDF = TF \times \log\frac{N}{DF}
  • Downweights common words

  • Improves BoW


3️⃣ Word Embeddings

Dense vectors that capture semantic meaning.

Popular methods:

  • Word2Vec

  • GloVe

  • FastText

Example:

king − man + woman ≈ queen

🔄 NLP Models

Traditional Models:

  • Naive Bayes

  • Logistic Regression

  • SVM

Deep Learning Models:

  • RNN

  • LSTM / GRU

  • CNN (for text)

  • Transformers (BERT, GPT)


🏷️ Common NLP Tasks

  • Text classification

  • Sentiment analysis

  • Named Entity Recognition (NER)

  • Part-of-Speech (POS) tagging

  • Machine translation

  • Question answering

  • Text summarization


🧠 Language Modeling

Predicting the next word in a sentence.

Example:

"I am learning deep ____" → learning

Used in:

  • Chatbots

  • Auto-complete

  • Text generation


📉 NLP Evaluation Metrics

  • Accuracy

  • Precision / Recall / F1-score

  • BLEU → translation

  • ROUGE → summarization

  • Perplexity → language models


🌍 Applications of NLP

  • Search engines

  • Chatbots & virtual assistants

  • Spam detection

  • Voice assistants

  • Social media analysis

  • Recommendation systems


🧠 NLP in One Line

NLP teaches machines to read, understand, and speak human language.


✅ Summary

  • NLP processes human language

  • Preprocessing cleans text

  • Vectorization converts words to numbers

  • Models learn patterns



🎯 What is Reinforcement Learning (RL)?

Reinforcement Learning is a type of machine learning where an agent learns to make decisions by interacting with an environment to maximize cumulative reward.

👉 No labeled data.
👉 Learning happens through trial and error.


🧠 Core Components of RL

An RL problem is defined by these elements:

ComponentMeaning
AgentLearner / decision-maker
EnvironmentWorld the agent interacts with
State (S)Current situation
Action (A)What the agent can do
Reward (R)Feedback from environment
Policy (π)Strategy for choosing actions

🔁 RL Interaction Loop

State (Sₜ) ↓ Agent selects Action (Aₜ) ↓ Environment gives Reward (Rₜ₊₁) ↓ Moves to next State (Sₜ₊₁)

This loop continues until a terminal state.


🧩 Markov Decision Process (MDP)

Most RL problems are modeled as an MDP.

An MDP is defined as:

(S,A,P,R,γ)(S, A, P, R, \gamma)

Where:

  • S = set of states

  • A = set of actions

  • P = transition probability

  • R = reward function

  • γ = discount factor


💰 Reward & Return

Immediate Reward

Reward received after an action.

Return (Cumulative Reward)

Gt=Rt+1+γRt+2+γ2Rt+3+...G_t = R_{t+1} + \gamma R_{t+2} + \gamma^2 R_{t+3} + ...
  • γ (0–1) controls importance of future rewards

  • High γ → long-term planning


🎯 Policy

A policy defines the agent’s behavior.

Types:

  • Deterministic: π(s) → a

  • Stochastic: π(a|s) → probability

Goal of RL:

Find an optimal policy π* that maximizes expected return.


📊 Value Functions

Used to estimate how good a state or action is.

State-Value Function

Vπ(s)=E[Gt∣St=s]V^\pi(s) = \mathbb{E}[G_t | S_t = s]

Action-Value Function (Q-function)

Qπ(s,a)=E[Gt∣St=s,At=a]Q^\pi(s, a) = \mathbb{E}[G_t | S_t = s, A_t = a]

🔍 Exploration vs Exploitation

  • Exploration → try new actions

  • Exploitation → use known best actions

Common strategy:

  • ε-greedy

    • With probability ε → explore

    • With probability 1−ε → exploit


🧠 Types of Reinforcement Learning

1️⃣ Model-Based RL

  • Agent knows environment model

  • Plans actions using transitions

2️⃣ Model-Free RL

  • Learns directly from experience

  • Most commonly used


🏆 Common RL Algorithms

Value-Based:

  • Q-Learning

  • SARSA

  • Deep Q-Networks (DQN)

Policy-Based:

  • REINFORCE

  • Policy Gradient

Actor-Critic:

  • A2C / A3C

  • PPO

  • DDPG


🧮 Q-Learning Update Rule

Q(s,a)←Q(s,a)+α[R+γmax⁡Q(s′,a′)−Q(s,a)]Q(s,a) \leftarrow Q(s,a) + \alpha [R + \gamma \max Q(s',a') - Q(s,a)]

Where:

  • α = learning rate

  • γ = discount factor


🎮 Applications of RL

  • Game playing (Chess, Go, Atari)

  • Robotics & control systems

  • Autonomous driving

  • Recommendation systems

  • Trading & portfolio optimization


🧠 RL in One Line

Learn by acting, evaluate by reward, improve by experience.


✅ Summary

  • RL learns via interaction

  • Rewards guide learning

  • Policies define actions

  • Value functions estimate goodness

  • Exploration balances learning



🤖 What are AI Frameworks?

AI frameworks are software libraries that make it easier to:

  • Build neural networks

  • Train models efficiently (GPU/TPU)

  • Handle automatic differentiation

  • Deploy models to production

Without frameworks → tons of manual math 😵
With frameworks → focus on ideas & experiments 🚀


🔷 TensorFlow

TensorFlow is an open-source deep learning framework developed by Google.

Key Features

  • Supports CPU, GPU, TPU

  • Strong production & deployment tools

  • Scalable for large systems

  • Integrates well with cloud (GCP)


TensorFlow Architecture

  • Tensor → multi-dimensional array

  • Graph-based computation (static graph)

  • Uses Keras as high-level API

Model → Compile → Train → Evaluate → Deploy

TensorFlow + Keras Example

import tensorflow as tf from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense model = Sequential([ Dense(64, activation='relu', input_shape=(10,)), Dense(1) ]) model.compile(optimizer='adam', loss='mse') model.fit(X_train, y_train, epochs=10)

Where TensorFlow Shines

  • Large-scale production systems

  • Mobile & web deployment (TF Lite, TF.js)

  • Industry-grade pipelines


🔶 PyTorch

PyTorch is an open-source framework developed by Meta (Facebook).

Key Features

  • Dynamic computation graph

  • Pythonic & intuitive

  • Easy debugging

  • Preferred in research & experimentation


PyTorch Philosophy

  • “Define-by-run”

  • Graph built on the fly

  • Feels like standard Python code


PyTorch Example

import torch import torch.nn as nn class Model(nn.Module): def __init__(self): super().__init__() self.fc = nn.Linear(10, 64) self.out = nn.Linear(64, 1) def forward(self, x): x = torch.relu(self.fc(x)) return self.out(x) model = Model()

⚖️ TensorFlow vs PyTorch

FeatureTensorFlowPyTorch
Computation GraphStaticDynamic
Ease of UseModerateVery easy
DebuggingHarderEasier
ResearchLess preferredHighly preferred
ProductionExcellentImproving fast

🔄 Automatic Differentiation

Both frameworks:

  • Automatically compute gradients

  • Use backpropagation

  • Support custom loss functions

Example (PyTorch):

loss.backward() optimizer.step()

🚀 Ecosystem & Tools

TensorFlow:

  • Keras

  • TensorFlow Lite

  • TensorFlow Serving

  • TensorFlow.js

PyTorch:

  • TorchVision

  • TorchText

  • TorchAudio

  • PyTorch Lightning


🧠 Which One Should You Learn?

👉 Beginner / Research / Experimentation → PyTorch
👉 Production / Deployment / Mobile → TensorFlow

🔥 Real-world tip: learn both, start with PyTorch.


📌 Industry Usage

  • TensorFlow → Google, Airbnb, Uber

  • PyTorch → Meta, Tesla, OpenAI, research labs


✅ Summary

  • AI frameworks simplify deep learning

  • TensorFlow excels in deployment

  • PyTorch excels in flexibility & research

  • Both support GPUs and auto-grad



🚀 What is Advanced Analytics?

Advanced Analytics goes beyond descriptive reports and dashboards to:

  • Predict future outcomes

  • Recommend optimal actions

  • React instantly to live data

It combines:

  • Statistics

  • Machine Learning / AI

  • Optimization

  • Streaming systems


📈 Analytics Evolution

Descriptive → Diagnostic → Predictive → Prescriptive → Real-Time

🔮 1. Predictive Analytics

“What is likely to happen?”

Predictive analytics uses historical data + ML models to forecast future events.

Techniques:

  • Regression (Linear, Logistic)

  • Time-series models (ARIMA, LSTM)

  • Classification models

  • Ensemble methods (Random Forest, XGBoost)

Examples:

  • Sales forecasting

  • Customer churn prediction

  • Stock price prediction

  • Credit risk assessment

Output:

  • Probabilities

  • Forecasted values


🧠 Predictive Analytics Workflow

Historical Data ↓ Feature Engineering ↓ Model Training ↓ Prediction ↓ Evaluation

🧩 2. Prescriptive Analytics

“What should we do?”

Prescriptive analytics goes beyond prediction and suggests optimal decisions.

Core Idea:

  • Combine predictions + constraints + objectives

  • Use optimization & simulation


Techniques:

  • Optimization (Linear / Integer Programming)

  • Decision Trees & Rules

  • Reinforcement Learning

  • Simulation & What-if analysis

Examples:

  • Dynamic pricing

  • Supply chain optimization

  • Portfolio allocation

  • Personalized recommendations

Output:

  • Action recommendations

  • Optimal strategies


⚙️ Predictive vs Prescriptive

AspectPredictivePrescriptive
FocusFuture outcomesBest actions
OutputPredictionRecommendation
TechniquesML modelsOptimization + ML
QuestionWhat will happen?What should we do?

⚡ 3. Real-Time Analytics

“What is happening right now?”

Real-time analytics processes live data streams and responds instantly.

Characteristics:

  • Low latency

  • Continuous processing

  • Event-driven decisions


Technologies:

  • Apache Kafka

  • Apache Spark Streaming

  • Apache Flink

  • AWS Kinesis

Examples:

  • Fraud detection

  • Stock trading systems

  • IoT monitoring

  • Recommendation updates


🧠 Real-Time Analytics Pipeline

Live Data ↓ Stream Ingestion ↓ Real-Time Processing ↓ ML Inference ↓ Action / Alert

🔁 How They Work Together

Real systems often combine all three:

Example (E-commerce):

  • Predictive → Forecast demand

  • Prescriptive → Optimize pricing

  • Real-Time → Adjust prices live


🏭 Industry Use Cases

  • Finance → fraud detection, algorithmic trading

  • Healthcare → patient risk monitoring

  • Manufacturing → predictive maintenance

  • Retail → personalization, inventory planning

  • Smart Cities → traffic optimization


🧠 Key Challenges

  • Data quality

  • Latency constraints

  • Model drift

  • Scalability

  • Explainability


🧠 One-Line Intuition

  • Predictive → foresee

  • Prescriptive → decide

  • Real-Time → react instantly


✅ Summary

  • Advanced analytics drives intelligent decisions

  • Predictive models forecast outcomes

  • Prescriptive analytics recommends actions

  • Real-time systems act immediately



📊 What is Big Data Analytics?

Big Data Analytics is the process of analyzing very large, fast, and complex datasets that traditional systems can’t handle.

The 5 V’s of Big Data:

  • Volume – massive data sizes (TBs, PBs)

  • Velocity – fast data generation (streams)

  • Variety – structured, semi-structured, unstructured

  • Veracity – data quality & uncertainty

  • Value – useful insights


🧠 Why Traditional Systems Fail

  • Limited storage

  • Single-machine processing

  • Slow batch execution

👉 Big data frameworks use distributed storage + parallel processing.


🐘 Apache Hadoop

Hadoop is an open-source framework for distributed storage and batch processing.

Core Components:

HDFS + YARN + MapReduce

1️⃣ HDFS (Hadoop Distributed File System)

  • Stores data across multiple machines

  • Splits files into blocks (default: 128MB)

  • Fault-tolerant via replication

Key nodes:

  • NameNode → metadata

  • DataNode → actual data


2️⃣ MapReduce

Programming model for batch processing.

Two phases:

  • Map → process & filter data

  • Reduce → aggregate results

Example:

  • Word Count

    • Map → (word, 1)

    • Reduce → (word, total count)


3️⃣ YARN

  • Resource manager

  • Allocates CPU & memory

  • Schedules jobs


🟡 Hadoop Pros & Cons

✅ Cheap storage
✅ Highly fault-tolerant
❌ Slow (disk-based)
❌ Complex programming


⚡ Apache Spark

Apache Spark is a fast, in-memory distributed data processing engine.

👉 Spark is ~100x faster than Hadoop MapReduce (for many workloads).


🔥 Spark Architecture

Driver Program ↓ Cluster Manager ↓ Executors

Key Spark Components:

  • Spark Core – basic processing

  • Spark SQL – structured data

  • Spark Streaming – real-time data

  • MLlib – machine learning

  • GraphX – graph processing


🧱 RDDs (Resilient Distributed Datasets)

  • Immutable distributed collections

  • Fault-tolerant

  • In-memory processing

Operations:

  • Transformations → lazy (map, filter)

  • Actions → execute (count, collect)


🧠 DataFrames & Datasets

Higher-level APIs than RDDs:

  • Optimized

  • Easier to use

  • SQL-like queries

df.filter(df.age > 25).groupBy("city").count()

🔄 Hadoop vs Spark

FeatureHadoopSpark
ProcessingDisk-basedIn-memory
SpeedSlowVery fast
Ease of useHardEasy
Real-time support❌✅
ML supportLimitedStrong

🔁 Hadoop + Spark Together

They are often used together:

  • Hadoop → storage (HDFS)

  • Spark → processing & analytics


🌍 Real-World Use Cases

  • Log analytics

  • Recommendation systems

  • Fraud detection

  • IoT data processing

  • Social media analytics


🧠 When to Use What?

  • Massive batch processing → Hadoop

  • Fast analytics & ML → Spark

  • Real-time streaming → Spark Streaming


✅ Summary

  • Big data analytics handles massive datasets

  • Hadoop provides distributed storage & batch processing

  • Spark provides fast, in-memory analytics

  • Spark is the modern big data engine



🧠 AI in Healthcare 🏥

Goal: Improve diagnosis, treatment, and patient care

Key Applications

  • Medical imaging

    • Tumor detection (CNNs on X-ray, MRI, CT)

  • Disease prediction

    • Diabetes, heart disease (ML models)

  • Drug discovery

    • Molecular modeling, protein folding

  • Personalized medicine

    • Treatment recommendations

  • Virtual health assistants

    • Symptom checking chatbots

AI Techniques Used

  • CNNs (image analysis)

  • NLP (clinical notes)

  • LSTM (patient time-series data)

  • Reinforcement Learning (treatment optimization)

Benefits

  • Early diagnosis

  • Reduced human error

  • Faster treatment decisions


💰 AI in Finance 📊

Goal: Risk reduction, automation, and profit optimization

Key Applications

  • Fraud detection

    • Real-time anomaly detection

  • Algorithmic trading

    • Reinforcement learning, LSTMs

  • Credit scoring

    • Loan default prediction

  • Customer segmentation

  • Robo-advisors

    • Automated investment strategies

AI Techniques Used

  • Supervised ML (classification)

  • Time-series models

  • Reinforcement Learning

  • Big data analytics (Spark)

Benefits

  • Faster decisions

  • Reduced fraud losses

  • Personalized financial services


🛒 AI in E-commerce

Goal: Personalization and conversion optimization

Key Applications

  • Recommendation systems

    • “You may also like”

  • Dynamic pricing

    • Demand-based pricing

  • Customer churn prediction

  • Chatbots & virtual assistants

  • Visual search

    • Search by image

AI Techniques Used

  • Collaborative filtering

  • NLP (chatbots, reviews)

  • CNNs (product images)

  • Predictive & prescriptive analytics

Benefits

  • Higher sales

  • Better user experience

  • Customer retention


🤖 AI in Robotics

Goal: Autonomous decision-making and control

Key Applications

  • Autonomous robots

    • Navigation & obstacle avoidance

  • Industrial robots

    • Assembly, quality inspection

  • Service robots

    • Delivery, cleaning, assistance

  • Humanoid robots

AI Techniques Used

  • Reinforcement Learning (control)

  • Computer Vision (CNNs)

  • Sensor fusion

  • SLAM (Simultaneous Localization and Mapping)

Benefits

  • Automation

  • Precision

  • Safety in hazardous environments


🔗 How AI Technologies Connect

DomainCore AI Tech
HealthcareCNN, NLP, LSTM
FinanceML, RL, Time-Series
E-commerceRecommender Systems, NLP
RoboticsRL, CV, Sensor Fusion

⚠️ Challenges Across Domains

  • Data privacy & security

  • Bias & fairness

  • Explainability (especially healthcare & finance)

  • High deployment cost

  • Regulatory constraints


🌍 Real-World Example (End-to-End)

E-commerce platform:

  • Big Data (Spark) → user behavior

  • Predictive Analytics → demand forecast

  • Prescriptive Analytics → pricing strategy

  • Real-Time AI → live recommendations


🧠 One-Line Takeaway

AI turns data into decisions, automation, and intelligence across industries.


✅ Summary

  • AI improves efficiency and accuracy

  • Healthcare → better diagnosis

  • Finance → risk & fraud control

  • E-commerce → personalization

  • Robotics → autonomy



🚀 What is AI Model Deployment?

Model deployment is the process of making a trained AI/ML model available for real-world use so it can generate predictions on new data.

Training a model ≠ using a model
Deployment = turning a model into a service or product


🧠 Typical ML Lifecycle

Data → Training → Evaluation → Deployment → Monitoring → Retraining

Deployment & monitoring are continuous, not one-time steps.


🏗️ Deployment Architectures

1️⃣ Batch Deployment

  • Predictions run on a schedule

  • Used for large datasets

Examples:

  • Monthly churn prediction

  • Daily sales forecasting

✅ Simple
❌ Not real-time


2️⃣ Real-Time (Online) Deployment

  • Model exposed as an API

  • Low-latency predictions

Examples:

  • Fraud detection

  • Recommendation systems

Common tools:

  • REST APIs (FastAPI, Flask)

  • Docker + Kubernetes


3️⃣ Edge Deployment

  • Model runs on local devices

  • No internet dependency

Examples:

  • Medical devices

  • Autonomous vehicles

  • Mobile apps

Tools:

  • TensorFlow Lite

  • ONNX

  • NVIDIA TensorRT


🧰 Common Deployment Tools

CategoryTools
APIFlask, FastAPI
ContainerDocker
OrchestrationKubernetes
CloudAWS SageMaker, GCP AI Platform
Model FormatONNX, SavedModel

🔄 CI/CD for ML (MLOps)

MLOps applies DevOps ideas to ML systems.

Key Components:

  • Version control (Git, DVC)

  • Automated testing

  • Continuous training

  • Automated deployment

Pipeline:

Code → Train → Test → Deploy → Monitor

📊 Model Monitoring

Monitoring ensures the model continues to perform well after deployment.


1️⃣ Data Drift

Input data changes over time.

  • Example: customer behavior shifts

Detection:

  • Statistical tests (KS-test)

  • Feature distribution monitoring


2️⃣ Concept Drift

Relationship between input & output changes.

  • Example: fraud patterns evolve

Harder to detect → needs performance tracking


3️⃣ Performance Monitoring

Track metrics like:

  • Accuracy

  • Precision / Recall

  • RMSE

  • Latency


🚨 Alerting & Logging

  • Log inputs, outputs, errors

  • Set thresholds for alerts

  • Monitor API failures

Tools:

  • Prometheus

  • Grafana

  • ELK Stack


🔁 Model Retraining Strategies

  • Scheduled retraining (weekly/monthly)

  • Drift-triggered retraining

  • Human-in-the-loop feedback


🔐 Security & Reliability

  • Model access control

  • Input validation

  • Adversarial attack protection

  • Rollback mechanisms


🌍 Real-World Example

Fraud Detection System

  1. Train model on historical data

  2. Deploy as REST API

  3. Monitor live transactions

  4. Detect drift

  5. Retrain model

  6. Redeploy seamlessly


⚠️ Common Challenges

  • Model decay over time

  • Data leakage

  • Scalability issues

  • Explainability requirements

  • Regulatory compliance


🧠 One-Line Insight

A model is only valuable if it works reliably in production.


✅ Summary

  • Deployment makes models usable

  • Monitoring keeps them reliable

  • MLOps automates the lifecycle

  • Retraining ensures long-term performance


⚖️ AI Ethics: What & Why

AI Ethics deals with building AI systems that are:

  • Fair

  • Transparent

  • Accountable

  • Safe

  • Respectful of human rights

Why it matters:

  • AI influences healthcare, finance, hiring, law

  • Poorly designed AI can cause real harm


🚨 Ethical Risks in AI

  • Discrimination & unfair decisions

  • Privacy violations

  • Lack of accountability

  • Automation bias (blind trust in AI)

  • Misuse & surveillance


🎭 Bias in AI

Bias occurs when AI systems produce systematically unfair outcomes.


🔍 Sources of Bias

1️⃣ Data Bias

  • Skewed or incomplete datasets

  • Historical inequalities

Example:

  • Hiring data biased toward one gender


2️⃣ Algorithmic Bias

  • Model design amplifies patterns

  • Optimization favors majority groups


3️⃣ Human Bias

  • Subjective labeling

  • Biased feature selection


⚠️ Types of Bias

  • Gender bias

  • Racial / ethnic bias

  • Age bias

  • Socioeconomic bias


🛠️ Bias Mitigation Strategies

  • Diverse & representative datasets

  • Bias-aware feature engineering

  • Fairness constraints in models

  • Regular audits & monitoring

Fairness metrics:

  • Demographic parity

  • Equal opportunity

  • Disparate impact


🔍 What is Explainable AI (XAI)?

Explainable AI (XAI) refers to techniques that make AI decisions understandable to humans.

Why XAI is needed:

  • Trust & transparency

  • Regulatory compliance

  • Debugging models

  • Ethical accountability


🧠 Black Box vs Glass Box

Model TypeExplainability
Linear RegressionHigh
Decision TreesHigh
Random ForestMedium
Deep Neural NetworksLow

🧩 XAI Techniques

1️⃣ Model-Intrinsic Methods

Explainable by design:

  • Linear models

  • Decision trees

  • Rule-based systems


2️⃣ Post-Hoc Explanation Methods

🔹 LIME

  • Explains individual predictions

  • Uses local approximations

🔹 SHAP

  • Based on game theory

  • Shows feature contribution

🔹 Feature Importance

  • Global model behavior

🔹 Saliency Maps (CNNs)

  • Highlight important image regions


🏥 XAI in High-Stakes Domains

  • Healthcare → diagnosis explanation

  • Finance → loan approval reasoning

  • Law → sentencing & risk scores

👉 Often legally required.


📜 Regulations & Guidelines

  • GDPR (Right to explanation)

  • AI governance frameworks

  • Model documentation (Model Cards)

  • Data Sheets for datasets


🧠 Ethical AI Principles (Quick List)

  • Fairness

  • Transparency

  • Accountability

  • Privacy

  • Human oversight


⚠️ Challenges in Ethical AI

  • Trade-off between accuracy & fairness

  • Explaining deep models

  • Cultural differences in ethics

  • Continuous monitoring


🧠 One-Line Takeaway

Ethical AI isn’t optional — it’s responsible engineering.


✅ Summary

  • Ethics ensures responsible AI use

  • Bias leads to unfair outcomes

  • XAI builds trust and accountability

  • Monitoring & governance are essential



🌟 Why AI Trends Matter

AI is moving toward:

  • Less manual effort

  • More automation

  • Human-like interaction

  • Faster deployment

These trends lower the barrier to entry and massively scale impact.


🤖 1. AutoML (Automated Machine Learning)

“AI that builds AI”

AutoML automates the end-to-end ML pipeline:

  • Data preprocessing

  • Feature engineering

  • Model selection

  • Hyperparameter tuning


How AutoML Works

Raw Data ↓ Auto Preprocessing ↓ Model Search ↓ Hyperparameter Optimization ↓ Best Model

Popular AutoML Tools

  • Google AutoML

  • H2O.ai

  • Auto-sklearn

  • TPOT

  • AWS SageMaker Autopilot


Use Cases

  • Rapid prototyping

  • Business analysts using ML

  • Baseline model creation


Pros & Cons

✅ Fast
✅ Reduces expertise barrier
❌ Limited customization
❌ Black-box risk


✨ 2. Generative AI

“AI that creates”

Generative AI produces new content, not just predictions.

Generates:

  • Text

  • Images

  • Audio

  • Code

  • Video


Key Models

  • Large Language Models (LLMs) – GPT, LLaMA

  • Diffusion models – image generation

  • GANs – realistic data synthesis

  • VAEs – probabilistic generation


How Generative AI Works (High Level)

  • Learns data distribution

  • Samples from learned space

  • Produces original but realistic outputs


Applications

  • Content creation

  • Code generation

  • Drug discovery

  • Synthetic data generation

  • Personalized education


Risks & Challenges

  • Hallucinations

  • IP & copyright issues

  • Bias amplification

  • Misuse (deepfakes)


💬 3. Chatbots & Conversational AI

“AI that talks”

Modern chatbots go far beyond rule-based systems.


Evolution of Chatbots

EraType
EarlyRule-based
MidML-based
NowLLM-powered

Core Components

  • NLP / NLU – understand intent

  • Dialogue management

  • Response generation

  • Context memory


Technologies Used

  • Transformers

  • LLMs

  • RAG (Retrieval-Augmented Generation)

  • Speech-to-Text & Text-to-Speech


Use Cases

  • Customer support

  • Virtual assistants

  • Healthcare triage

  • HR & IT helpdesks

  • Education tutors


🔗 How These Trends Connect

Example (Business AI System):

  • AutoML → build prediction models

  • Generative AI → generate insights & reports

  • Chatbots → deliver insights conversationally


🧠 Impact on Analytics

  • Shift from dashboards → conversations

  • From manual modeling → automated pipelines

  • From static reports → generated insights


🔮 Future Directions

  • Multi-modal AI (text + image + audio)

  • AI agents (task-performing systems)

  • Stronger AI governance

  • Human-AI collaboration


🧠 One-Line Insight

AI is becoming more automated, more creative, and more conversational.


✅ Summary

  • AutoML democratizes ML

  • Generative AI creates content

  • Chatbots enable natural interaction

  • Together, they redefine analytics & AI systems











4 comments:

02

Capstone resource hub

Codingacharya

Capstone Learning Resources, Notes & Project Hub

TCS NQT Questions
Read Notes
Machine Learning – ACE Theory
Read Notes
Machine Learning PPT
Read Notes
MachienLearning LAB
Read Notes
CSPT LAB programs
Read Notes
Time table and CSPT syllabus
Read Notes
Appreciations
Read Notes
ISTE life memberships
Read Notes
Artificial Intelligence & Analytics
Read Notes
Fullstack Web Dev
Read Notes
MERN Web Dev
Read Notes
Course Structure
Read Notes
Cloud Computing
Read Notes
90 Days ML Challenge
Read Notes
Advanced Analytics & Viz
Read Notes
Advanced Machine Learning
Read Notes
React JS
Read Notes
ML Chaitanya
Read Notes
Important Links
Read Notes
CSS Effects
Read Notes
RESUME
Read Notes
Bootstrap CSS
Read Notes
MongoDB
Read Notes
OWN Python Package
Read Notes
HTML Course
Read Notes
HTML Projects
Read Notes
GitHub Projects
Read Notes
Angular JS
Read Notes
Journals
Read Notes
NLP Notes
Read Notes
Videos
Read Notes
Data Analytics & Viz
Read Notes
Cloud Computing (Archive)
Read Notes
Open CV
Read Notes
jQuery
Read Notes
React JS (Archive)
Read Notes
Node JS
Read Notes
DAV Theory
Read Notes
DAV Lab
Read Notes
Big Data Notes
Read Notes
R-Programming
Read Notes
HADOOP Lab
Read Notes
GATE DA
Read Notes
JAVA Lab
Read Notes
Computer Networks
Read Notes
03

Live projects & profiles