Data science is the field that combines statistics,
programming, and domain knowledge to extract useful insights
and knowledge from data.
At its core, data science involves:
1.
Collecting
data – from databases, sensors, APIs, websites, or other
sources.
2.
Cleaning
and processing data – handling missing values, noise, and
formatting issues.
3.
Exploring
and analyzing data – using statistics and visualization to
understand patterns and trends.
4.
Building
models – applying machine learning (ML) and artificial
intelligence (AI) to make predictions, detect anomalies, or classify data.
5.
Communicating
results – presenting insights with dashboards, reports, or visualizations
for decision-making.
It’s an interdisciplinary
field that sits at the intersection of:
·
Mathematics
& Statistics (probability, hypothesis testing, regression,
etc.)
·
Computer
Science (Python, R, SQL, big data tools, ML libraries)
·
Domain
Expertise (finance, healthcare, marketing, engineering, etc.)
In simple words: Data science is about turning raw data into actionable insights.
Data Science Life Cycle
The Data Science Life Cycle
describes the step-by-step process data scientists follow to turn raw data into
meaningful insights or predictive models. It usually includes these stages:
1.
Problem Definition
- Understand the business or research problem.
- Define the objective (e.g., “Predict customer churn” or
“Classify tumors as benign or malignant”).
2.
Data Collection
- Gather data from various sources: databases, APIs,
sensors, web scraping, surveys, etc.
- Ensure data relevance and availability.
3.
Data Cleaning & Preparation
- Handle missing values, duplicates, and outliers.
- Transform raw data into structured formats.
- Feature engineering (create new variables).
- Normalize, scale, or encode categorical data.
4.
Exploratory Data Analysis (EDA)
- Use statistical methods and visualization to identify
trends, correlations, and patterns.
- Example tools: Pandas, Matplotlib, Seaborn, Tableau.
5.
Modeling
- Choose suitable machine learning/statistical models
(Regression, Decision Trees, Neural Networks, etc.).
- Train models on training data.
- Optimize hyperparameters.
6.
Evaluation
- Test the model with validation/test datasets.
- Use metrics (Accuracy, Precision, Recall, F1-score,
RMSE, R², etc.) depending on the problem type.
7.
Deployment
- Integrate the model into production (web app, API,
dashboard, etc.).
- Example: A recommendation engine on Netflix or fraud
detection system in banking.
8.
Monitoring & Maintenance
- Continuously track performance (data drift, model
accuracy).
- Retrain/update models with new data.
Data Scientist
Focus:
Extract insights and build predictive models from data.
Key Responsibilities:
·
Understand business problems and translate them
into data problems.
·
Collect, clean, and analyze large datasets.
·
Build statistical and machine learning models.
·
Communicate insights with visualizations and
reports.
·
Prototype ML models but may not always deploy
them.
Skills:
·
Python/R, SQL, Pandas, NumPy, Scikit-learn.
·
Statistics, probability, hypothesis testing.
·
Data visualization (Matplotlib, Seaborn,
Tableau, Power BI).
·
Basic ML & sometimes deep learning.
Example Work: Predict customer churn, recommend products, detect fraud.
Data Analyst
Focus:
Analyze historical data to provide actionable insights.
Key Responsibilities:
·
Collect, clean, and organize data.
·
Perform exploratory data analysis (EDA).
·
Create dashboards and reports.
·
Answer business questions using SQL and BI
tools.
·
Focus more on descriptive & diagnostic analytics
(what happened, why it happened).
Skills:
·
SQL (very strong).
·
Excel, Tableau, Power BI.
·
Python/R (sometimes, for advanced analysis).
·
Strong business communication.
Example Work: Sales trend analysis, KPI reports, customer behavior dashboards.
Machine Learning Engineer
Focus:
Deploy and scale machine learning models in production.
Key Responsibilities:
·
Take ML prototypes (often from data scientists)
and make them production-ready.
·
Build data pipelines and model deployment
workflows.
·
Optimize performance, scalability, and latency
of models.
·
Work with software engineers and cloud
infrastructure (AWS, GCP, Azure).
·
Focus more on engineering and automation.
Skills:
·
Strong programming (Python, Java, C++).
·
ML frameworks (TensorFlow, PyTorch,
Scikit-learn).
·
MLOps (Docker, Kubernetes, CI/CD).
·
Cloud platforms (AWS SageMaker, GCP AI
Platform).
·
Data pipelines (Airflow, Spark).
Example Work: Deploying a real-time recommendation engine, fraud detection API, or autonomous vehicle perception system.
Quick Comparison Table
|
Role |
Main Goal |
Tools/Skills |
Example Output |
|
Data Analyst |
Understand what happened |
SQL, Excel, Tableau, BI tools |
Dashboard/report |
|
Data Scientist |
Predict what will happen |
Python/R, ML, Stats, Viz |
Predictive model, insights |
|
ML Engineer |
Make models work in real-time |
Python, TensorFlow, MLOps |
Deployed ML system/API |
In short:
·
Data
Analyst = Insights & reporting.
·
Data
Scientist = Insights + ML modeling.
· ML Engineer = Productionizing ML models.
Types of Data
1. By Nature (Qualitative vs Quantitative)
Qualitative Data
(Categorical)
·
Describes qualities or characteristics
(non-numerical).
·
Examples: Gender (Male/Female), Colors (Red,
Blue), City (Paris, Delhi).
Types of Qualitative
Data:
·
Nominal
→ Categories without order (e.g., Blood group: A, B, AB, O).
·
Ordinal
→ Categories with order (e.g., Education level: High school < Bachelor <
Master < PhD).
Quantitative Data
(Numerical)
·
Describes measurable quantities.
·
Examples: Age, Salary, Temperature, Height.
Types of Quantitative
Data:
·
Discrete
→ Countable (e.g., Number of students = 30).
· Continuous → Measurable (e.g., Weight = 65.3 kg, Time = 12.5 sec).
2. By Structure
·
Structured
Data → Organized in rows/columns (e.g., Databases, Excel
sheets).
·
Unstructured
Data → No fixed format (e.g., Images, Videos, Text, Social
media posts).
· Semi-structured Data → Partly organized (e.g., JSON, XML, NoSQL databases).
3. By Time Dependency
·
Static
Data → Doesn’t change over time (e.g., Census data).
·
Dynamic
Data → Continuously updated (e.g., Stock prices, Sensor data).
· Real-time / Streaming Data → Generated instantly (e.g., IoT sensors, live tweets, online transactions).
4. By Source
·
First-party
data → Collected directly by an organization (e.g., Customer
surveys, website analytics).
·
Second-party
data → Shared by another company (partnership data).
· Third-party data → Purchased from external providers (e.g., Market research data).
Summary:
·
Qualitative
vs Quantitative (what type of values?)
·
Structured
vs Unstructured (how is it organized?)
·
Static
vs Dynamic (how does it change over time?)
· First/Second/Third-party (where does it come from?)
1. Primary Data Collection (Direct from Source)
Collected first-hand by researchers or
organizations.
·
Surveys
& Questionnaires 📝
o Used
for customer feedback, market research, opinions.
o Tools:
Google Forms, Typeform, SurveyMonkey.
·
Interviews
🎤
o One-on-one
discussions to gather detailed insights.
·
Observation
👀
o Watching
and recording behaviors (e.g., traffic monitoring, store visits).
·
Experiments
🔬
o Controlled
setups (e.g., A/B testing, clinical trials).
·
Sensors/IoT
Devices 📡
o Data from machines, GPS, smart devices, health trackers.
2. Secondary Data Collection
(Already Available Data)
Using pre-existing sources of data.
·
Databases
& Data Warehouses
o SQL/NoSQL
databases, company ERP/CRM systems.
·
Public
Datasets
o Kaggle,
UCI Machine Learning Repository, Government portals (data.gov).
·
APIs
o Twitter
API, Google Maps API, Financial market APIs.
·
Web
Scraping 🌐
o Extracting
data from websites using Python (BeautifulSoup, Scrapy, Selenium).
·
Reports
& Research Papers
o Published studies, industry reports, academic journals.
3. Automated Data Collection
·
Tracking
Tools
o Google
Analytics, Mixpanel, Hotjar for user behavior.
·
Logs
o Server
logs, clickstream data.
·
Streaming
Data
o Real-time feeds from Kafka, Spark Streaming, IoT pipelines.
4. Crowd-sourced &
Collaborative Data
·
Wikipedia contributions.
·
OpenStreetMap.
· GitHub datasets & community-driven data.
Summary:
·
Primary:
Surveys, interviews, experiments, sensors.
·
Secondary:
Databases, APIs, public datasets, web scraping.
·
Automated:
Logs, analytics tools, streaming data.
· Crowdsourced: Community-generated data.
data cleaning and pre processing
ChatGPT
said:
Great question 👍 — Data Cleaning and Preprocessing is one of the most
important steps in the data science life cycle.
(There’s a saying: “80% of a data
scientist’s time is spent cleaning data, 20% modeling it” 😅).
🧹 Data Cleaning
The process of fixing or removing incorrect,
incomplete, irrelevant, or duplicated data.
🔹 Common Steps:
1.
Handling Missing
Data
o Remove
rows/columns with too many missing values.
o Impute
missing values (mean, median, mode, interpolation, ML-based imputation).
2.
Removing Duplicates
o Drop
duplicate rows/records.
3.
Correcting Data
Types
o Convert
strings to dates, numbers to categorical, etc.
4.
Handling Outliers
o Detect
with Z-score, IQR, or visualization (boxplot).
o Decide
whether to remove, cap, or transform.
5.
Fixing Inconsistent
Data
o Standardize
formats (e.g., “NY”, “New York” → “New York”).
o Handle typos and case sensitivity.
Data Preprocessing
The process of transforming raw data into a
format suitable for analysis/modeling.
Common Steps:
1.
Data
Normalization/Standardization
o Normalization (Min-Max scaling): Scales
values between 0–1.
o Standardization (Z-score scaling): Mean
= 0, SD = 1.
o Example:
Feature scaling before ML models (KNN, SVM).
2.
Encoding
Categorical Variables
o Label Encoding: Assigns numbers (Male=0,
Female=1).
o One-Hot Encoding: Creates binary columns
(Red=[1,0,0], Blue=[0,1,0]).
3.
Feature
Engineering
o Creating
new features (e.g., extracting "Year" from a date).
o Combining
or splitting columns.
4.
Feature
Selection/Dimensionality Reduction
o Remove
irrelevant or highly correlated features.
o Use
PCA, LDA, or feature importance methods.
5.
Data
Transformation
o Log
transformation, polynomial features, binning continuous variables.
6.
Train-Test Split
o Divide
data into train, validation, test sets
before modeling.
🟡 1. Handling Missing Values
Missing data can distort analysis and model accuracy.
🔹 Methods to Handle:
-
Remove Missing Data
-
Drop rows/columns with too many missing values.
-
Works if dataset is large and missing data is small.
-
-
Imputation (Fill Missing Values)
-
Mean/Median/Mode Imputation
-
Forward/Backward Fill
-
Interpolation
-
Model-Based Imputation (e.g., using KNN or regression to predict missing values).
-
🔴 2. Handling Outliers
Outliers = extreme values that deviate from the majority of data.
They can skew results or sometimes represent important rare events (like fraud detection).
🔹 Methods to Detect Outliers:
-
Statistical Methods
-
Z-Score Method (values with |z| > 3 are outliers).
-
IQR (Interquartile Range) Method
-
-
Visualization
-
Boxplot, Histogram, Scatterplot help detect extreme values.
-
-
Domain Knowledge
-
Example: A height of 300 cm is unrealistic (error). But in finance, extreme values (stock crashes) may be valid and important.
-
🔹 Ways to Handle Outliers:
-
Remove them (if they are errors/noise).
-
Cap them (Winsorization → replace extreme values with nearest boundary).
-
Transform them (log, square root to reduce skewness).
-
Use robust models (e.g., tree-based models handle outliers better than linear regression).
🟢 Why Encoding is Needed?
Categorical features like "Male/Female", "Red/Blue/Green", "Yes/No" must be converted into numbers so algorithms can process them.
🔹 Types of Encoding Methods
1️⃣ Label Encoding
-
Assigns a unique integer to each category.
-
Simple but introduces ordinal relationships (which may not exist).
👉 Best for ordinal data (e.g., "Low < Medium < High").
2️⃣ One-Hot Encoding
-
Creates binary columns for each category.
-
Avoids artificial order but increases dimensionality.
👉 Best for nominal data (unordered categories like city names).
3️⃣ Ordinal Encoding
-
Map categories to numbers based on their order.
👉 Best for features with clear ranking.
4️⃣ Target / Mean Encoding
-
Replace category with the mean of target variable for that category.
👉 Useful for high-cardinality categorical data (e.g., thousands of ZIP codes).
⚠️ Risk of data leakage, so apply only on training set (with cross-validation).
5️⃣ Frequency / Count Encoding
-
Replace each category with its frequency or count.
👉 Simple and effective for large categorical features.
6️⃣ Binary Encoding (Advanced)
-
Converts categories into binary digits → reduces dimensionality.
-
Example: Category 1 = 001, Category 2 = 010, etc.
-
Libraries:
category_encoders
✅ Summary Table
| Encoding Method | Use Case | Pros | Cons |
|---|---|---|---|
| Label Encoding | Ordinal data | Simple | Imposes order (not for nominal data) |
| One-Hot | Nominal data, few categories | No false order | High dimensionality |
| Ordinal Encoding | Ordered categories | Preserves ranking | Assumes correct order |
| Target/Mean | High cardinality | Captures target relationship | Risk of leakage |
| Frequency/Count | Large categorical vars | Simple, compact | Ignores target variable |
| Binary Encoding | High-cardinality features | Reduces dimensions | Harder to interpret |
🔹 Steps in EDA
1️⃣ Understand the Dataset
-
Check dataset size, structure, datatypes.
2️⃣ Check Missing & Duplicate Data
3️⃣ Univariate Analysis (one variable at a time)
-
Categorical Variables → bar plots, value counts.
-
Numerical Variables → histograms, boxplots, density plots.
4️⃣ Bivariate Analysis (relationship between two variables)
-
Categorical vs Numerical → boxplots, violin plots.
-
Numerical vs Numerical → scatter plots, correlation.
5️⃣ Multivariate Analysis
-
Correlation heatmaps.
-
Pairplots for multiple relationships.
6️⃣ Outlier Detection
-
Use boxplots, histograms, Z-score, IQR.
7️⃣ Feature Engineering Ideas
-
Extract new features (e.g., Year from Date).
-
Encode categorical variables.
-
Transform skewed data (log, sqrt).
📊 EDA Tools
-
Python: Pandas, Matplotlib, Seaborn, Plotly.
-
Automated EDA Libraries:
pandas-profiling,sweetviz,dtale,autoviz. -
BI Tools: Power BI, Tableau for interactive dashboards.
✅ Summary
EDA helps you:
-
Understand data distributions.
-
Detect missing values & outliers.
-
Identify relationships between variables.
-
Guide feature engineering & model selection.
🎨 Matplotlib
-
Low-level visualization library in Python.
-
Very flexible → you can control almost every aspect of the plot (axes, ticks, colors).
-
But code can get verbose for complex plots.
Example:
🌈 Seaborn
-
Built on top of Matplotlib → simpler and prettier by default.
-
High-level → designed for statistical data visualization.
-
Integrates well with Pandas DataFrames.
-
Comes with built-in themes and advanced plots (heatmap, pairplot, violinplot, etc.).
Example:
🔹 Comparison Table
| Feature | Matplotlib 🖊️ | Seaborn 🎨 |
|---|---|---|
| Level | Low-level (more control) | High-level (easier, faster) |
| Ease of Use | Verbose | Concise, built-in themes |
| Data Handling | Works with lists/arrays | Works well with DataFrames |
| Plot Types | Basic (line, bar, hist) | Advanced (heatmap, pairplot, violin) |
| Customization | Full control | Limited (but can use Matplotlib for tweaks) |
✅ When to Use What?
-
Use Matplotlib → when you need full customization (scientific plots, journals).
-
Use Seaborn → when you need quick, beautiful statistical plots for data analysis.
-
Often → use them together (Seaborn for quick plots, Matplotlib to fine-tune).
No comments:
Post a Comment