student notes / est. for the classroom

HTML, CSS, JavaScript, Python, data science, computer networks — written the way you'd explain it to a classmate, not a compiler.

Top Job & Internship Portals

Handpicked portals for fresher jobs, tech roles, and listings in Hyderabad

GFG

GeeksforGeeks

Tech & Software Roles

Visit →
INT

Internshala

Fresher Jobs & Internships

Visit →
GOOG

Google Careers

Global Google Openings

Visit →
APN

Apna Jobs

Local Jobs in Hyderabad

Visit →
INS

Instahyre

Tech Roles in Hyderabad

Visit →
NAUK

Naukri.com

Fresher Jobs in Hyderabad

Visit →
📢 Updated daily

Internship & Job Alerts

01

Latest notes

February 23, 2024

BIG DATA Notes

 

Data is a set of characters used to collect, store and transmit information for a specific purpose. Data can be in any form, i.e., text, image, audio, etc. Data comes from the Latin word 'Datum', which means 'something given'. When the data is processed, it is termed as 'Information'.

 



What is Big Data?

Big Data refers to a collection of a very large and complicated set of data for which it becomes difficult to process using traditional or manual database management tools. The size of data grows exponentially and is generally in terabytes or more.

For example- over 500 million Tweets are generated on Twitter daily; Netflix has over 220 million paid memberships globally; there are over 2 billion daily users of Facebook. These statistics are quite large in numbers and increasing exponentially every year and thus can be classified as Big Data.

How Do We Classify Any Data as Big Data?

To classify any data set as Big Data, 3V's of Big Data were introduced in 2001, which later got updated to 5V's. These 5V's are:

1.      Volume: Volume refers to the 'size'or amount of data. For instance, YouTube has over 2.6 billion monthly active users and generates a large amount of data daily, which can't be processed manually; thus, modern techniques and tools are used to handle such voluminous data.

2.      Velocity: Velocity refers to the 'speed'or rate with which the data is accumulated. In 2010, YouTube had 200 million monthly active users, which increased to 2.6 billion in 2022.

3.      Variety: Variety refers to the 'heterogeneity' or diversity of data. The data can be structured, unstructured, or semi-structured.

4.      Veracity: Veracity refers to the 'trustworthiness'or quality of data. It means whether the data is free from various ambiguities or not.

5.      Value: Value refers to the 'Insights' gained from the data. It means whether the given data set is producing any useful result. Data, in its raw form, gives no valuable result, but once processed efficiently, it can give us important insights that could help us in decision-making.

Types of Big Data

There are three types of Big Data: Structured, Semi-structured and Unstructured data.

1.      Structured Data: Any data in a fixed format is known as structured data. It can only be accessed, stored, or processed in a particular format. This type of data is stored in the form of tables with rows and columns. Any Excel file or SQL file is an example of structured data.

2.      Unstructured Data: Unstructured data do not have a fixed format. These are stored in an unknown format. Such type of data is known as unstructured data. An example of unstructured data is a web page with text, images, videos, etc.

3.      Semi-structured Data: Semi-structured data is the combination of structured as well as unstructured forms of data. It does not contain any table to show relations; it contains tags or other markers to show hierarchy. JSON files, XML files, and CSV files (Comma-separated files) are semi-structured data examples. The e-mails we send or receive are also an example of semi-structured data.


Use Cases of Big Data

1.      Social Media and Entertainment: You must have witnessed streaming service apps such as Netflix recommending shows and movies based on your previous searches and what you have watched. It is done using the concept of Big Data. Netflix and other streaming service apps create a custom user profile, where they store the data of users, including their search history, their history, which genre they watch the most, at what time of day they prefer to watch the most, their streaming time per day, etc. analyze it and accordingly gives recommendations. It helps in a better streaming experience for the users.

2.      Shopping: Websites like Amazon, Flipkart, etc., also use Big Data to recommend products based on your previous purchases, search history, and interests. It is done to maximize their profits and provide a better shopping experience to their customers.

3.      Education: Big Data helps in analyzing and monitoring the behavior and activities of students, like the time they need to answer a question, the number of questions skipped, and the difficulty level of the questions that are skipped, and thus helps students to analyze their overall preparation, weak topics, strong topics, etc.

4.      Healthcare: Healthcare sectors use Big Data to track and analyze the health and fitness of the patients, the number of visits, the number of skipped appointments a patient, etc. Mass outbreaks of diseases can be predicted by analyzing the data and using algorithms.

5.      Transportation: Traffic control by collecting and analyzing the data from several sensors and cameras installed on roads and highways. Accident-prone areas can be detected with the help of Big Data analysis; thus, required measures can be taken to avoid accidents.

Evolution of Big Data

  • The earliest record to track and analyze data was not decades back but thousands of years back when accounting was first introduced in Mesopotamia.
  • In the 20th century, IBM developed the first large-scale data project, punch carding systems, which tracked the information of millions of Americans.
  • With the emergence of the World Wide Web and supercomputers in the 1990s, the creation of data on a large scale started to grow at an exponential rate. It was in the early 1990s when the term 'Big Data' was first used.
  • The two main challenges regarding 'Big Data' were storing and processing such a huge volume of data.
  • In 2005, Yahoo created the open-source framework Hadoop, which stores and processes large data sets.
  • The storage solution in Hadoop was named HDFS (Hadoop Distributed File System), and the processing solution was named MapReduce.
  • Later, Hadoop was handed over to an open-source and non-profitable corporation: Apache Software Foundation.
  • In 2008, Cloudera became the first company to provide commercial Hadoop distribution.
  • In 2013, the Creators of Apache Spark founded a company, Databricks, which offers a platform for Big Data and Machine Learning solutions.
  • Over the past few years, top Cloud providers such as Microsoft, Google, and Amazon also started to provide Big Data solutions. These Cloud providers made it much easier for users and companies to work on Big Data.

Did You Know?

In 2009, the Indian government stored fingerprints and iris scans of all its citizens in the largest database ever created.

A Brief Introduction to Hadoop

Founded by Doug Cutting and Mike Cafarella in 2005, Hadoop is an open-source framework that efficiently stores and processes Big Data. Hadoop is a Java-based framework. Apache Software Foundation manages Hadoop. The main components of Hadoop are HDFS (Hadoop Distributed File System) & MapReduce. Being an open-source platform, Hadoop is cost-efficient. Its speed and capacity to store large volumes of data make it popular among many top-tier companies. Companies such as Facebook, Twitter, LinkedIn, etc., use Hadoop to handle Big Data.

Importance of Big Data

1.      A better understanding of market conditions.

2.      Time and cost saving.

3.      Solving advertisers' problems.

4.      Offering better market insights.

5.      Boosting customer acquisition and retention.

Applications of Big Data

Big Data finds applications in various sectors, such as-

1.      Banking and Security

2.      Social Media and Entertainment

3.      E-commerce websites

4.      HealthCare

5.      Education

6.      Transportations

Big Data Analytics

Big Data Analytics uses modern tools and techniques to extract valuable insights, trends, hidden patterns, and relations with the help of large sets of data, which can be structured, semi-structured, or unstructured. It helps in better decision-making and optimizes business operations.

Let's consider the example of YouTube, which has over 2.6 billion monthly active users. It generates a huge amount of data every day. With the help of this data, it recommends videos based on what you have watched previously, your likes, shares, etc. What enables this is the tools and frameworks resulting from Big Data Analytics.

Types of Big Data Analytics

1.      Descriptive Analytics: This type of analytics summarizes or extracts insights based on the incoming We came up with a description based on the data. For example, insights drawn for your YouTube channel are based on the data such as likes, shares, and views on your videos.

2.      Predictive Analytics: This type of analytics predicts what might happen. Questions such as 'how' and 'why' reveal particular patterns that help predict future trends. Machine Learning concepts are also used for such types of analysis. For example, prediction of weather, prediction of malfunctioning in the parts of an airplane, etc.

3.      Prescriptive Analytics: These types of analytics are based on rules and recommendations and thus prescribe an analytical path. The analysis is generally based on the question, 'what actions should be taken?' Google's self-driving car is an example of prescriptive analysis.

4.      Diagnostic Analytics: These analytics look into past trends and diagnose questions such as how and why something happened. It is also called behavioral analytics. This analysis aims to answer the question, 'why did this happen?' For example, if the sales report of a company shows a rise in sales, then the company can analyze the internal and external causes responsible for the increase.

 

 

What is Hadoop

Hadoop is an open source framework from Apache and is used to store process and analyze data which are very huge in volume. Hadoop is written in Java and is not OLAP (online analytical processing). It is used for batch/offline processing.It is being used by Facebook, Yahoo, Google, Twitter, LinkedIn and many more. Moreover it can be scaled up just by adding nodes in the cluster.

Modules of Hadoop

1.      HDFS: Hadoop Distributed File System. Google published its paper GFS and on the basis of that HDFS was developed. It states that the files will be broken into blocks and stored in nodes over the distributed architecture.

2.      Yarn: Yet another Resource Negotiator is used for job scheduling and manage the cluster.

3.      Map Reduce: This is a framework which helps Java programs to do the parallel computation on data using key value pair. The Map task takes input data and converts it into a data set which can be computed in Key value pair. The output of Map task is consumed by reduce task and then the out of reducer gives the desired result.

4.      Hadoop Common: These Java libraries are used to start Hadoop and are used by other Hadoop modules.

Hadoop Architecture

The Hadoop architecture is a package of the file system, MapReduce engine and the HDFS (Hadoop Distributed File System). The MapReduce engine can be MapReduce/MR1 or YARN/MR2.

A Hadoop cluster consists of a single master and multiple slave nodes. The master node includes Job Tracker, Task Tracker, NameNode, and DataNode whereas the slave node includes DataNode and TaskTracker.


Hadoop Distributed File System

The Hadoop Distributed File System (HDFS) is a distributed file system for Hadoop. It contains a master/slave architecture. This architecture consist of a single NameNode performs the role of master, and multiple DataNodes performs the role of a slave.

Both NameNode and DataNode are capable enough to run on commodity machines. The Java language is used to develop HDFS. So any machine that supports Java language can easily run the NameNode and DataNode software.

NameNode

  • It is a single master server exist in the HDFS cluster.
  • As it is a single node, it may become the reason of single point failure.
  • It manages the file system namespace by executing an operation like the opening, renaming and closing the files.
  • It simplifies the architecture of the system.

DataNode

  • The HDFS cluster contains multiple DataNodes.
  • Each DataNode contains multiple data blocks.
  • These data blocks are used to store data.
  • It is the responsibility of DataNode to read and write requests from the file system's clients.
  • It performs block creation, deletion, and replication upon instruction from the NameNode.

AD

 Job Tracker

  • The role of Job Tracker is to accept the MapReduce jobs from client and process the data by using NameNode.
  • In response, NameNode provides metadata to Job Tracker.

Task Tracker

  • It works as a slave node for Job Tracker.
  • It receives task and code from Job Tracker and applies that code on the file. This process can also be called as a Mapper.

AD

 MapReduce Layer

The MapReduce comes into existence when the client application submits the MapReduce job to Job Tracker. In response, the Job Tracker sends the request to the appropriate Task Trackers. Sometimes, the TaskTracker fails or time out. In such a case, that part of the job is rescheduled.

Advantages of Hadoop

  • Fast: In HDFS the data distributed over the cluster and are mapped which helps in faster retrieval. Even the tools to process the data are often on the same servers, thus reducing the processing time. It is able to process terabytes of data in minutes and Peta bytes in hours.
  • Scalable: Hadoop cluster can be extended by just adding nodes in the cluster.
  • Cost Effective: Hadoop is open source and uses commodity hardware to store data so it really cost effective as compared to traditional relational database management system.
  • Resilient to failure: HDFS has the property with which it can replicate data over the network, so if one node is down or some other network failure happens, then Hadoop takes the other copy of data and use it. Normally, data are replicated thrice but the replication factor is configurable.

History of Hadoop

The Hadoop was started by Doug Cutting and Mike Cafarella in 2002. Its origin was the Google File System paper, published by Google.


Let's focus on the history of Hadoop in the following steps: -

  • In 2002, Doug Cutting and Mike Cafarella started to work on a project, Apache Nutch. It is an open source web crawler software project.
  • While working on Apache Nutch, they were dealing with big data. To store that data they have to spend a lot of costs which becomes the consequence of that project. This problem becomes one of the important reason for the emergence of Hadoop.
  • In 2003, Google introduced a file system known as GFS (Google file system). It is a proprietary distributed file system developed to provide efficient access to data.
  • In 2004, Google released a white paper on Map Reduce. This technique simplifies the data processing on large clusters.
  • In 2005, Doug Cutting and Mike Cafarella introduced a new file system known as NDFS (Nutch Distributed File System). This file system also includes Map reduce.
  • In 2006, Doug Cutting quit Google and joined Yahoo. On the basis of the Nutch project, Dough Cutting introduces a new project Hadoop with a file system known as HDFS (Hadoop Distributed File System). Hadoop first version 0.1.0 released in this year.
  • Doug Cutting gave named his project Hadoop after his son's toy elephant.
  • In 2007, Yahoo runs two clusters of 1000 machines.
  • In 2008, Hadoop became the fastest system to sort 1 terabyte of data on a 900 node cluster within 209 seconds.

Year

Event

2003

Google released the paper, Google File System (GFS).

2004

Google released a white paper on Map Reduce.

2006

  • Hadoop introduced.
  • Hadoop 0.1.0 released.
  • Yahoo deploys 300 machines and within this year reaches 600 machines.

2007

  • Yahoo runs 2 clusters of 1000 machines.
  • Hadoop includes HBase.

2008

  • YARN JIRA opened
  • Hadoop becomes the fastest system to sort 1 terabyte of data on a 900 node cluster within 209 seconds.
  • Yahoo clusters loaded with 10 terabytes per day.
  • Cloudera was founded as a Hadoop distributor.

2009

  • Yahoo runs 17 clusters of 24,000 machines.
  • Hadoop becomes capable enough to sort a petabyte.
  • MapReduce and HDFS become separate subproject.

2010

  • Hadoop added the support for Kerberos.
  • Hadoop operates 4,000 nodes with 40 petabytes.
  • Apache Hive and Pig released.

2011

  • Apache Zookeeper released.
  • Yahoo has 42,000 Hadoop nodes and hundreds of petabytes of storage.

2012

Apache Hadoop 1.0 version released.

2013

Apache Hadoop 2.2 version released.

2014

Apache Hadoop 2.6 version released.

2015

Apache Hadoop 2.7 version released.

2017

Apache Hadoop 3.0 version released.

2018

Apache Hadoop 3.1 version released.

  • In 2013, Hadoop 2.2 was released.
  • In 2017, Hadoop 3.0 was released.

What is Hadoop Streaming?

It is a utility or feature that comes with a Hadoop distribution that allows developers or programmers to write the Map-Reduce program using different programming languages like Ruby, Perl, Python, C++, etc. We can use any language that can read from the standard input(STDIN) like keyboard input and all and write using standard output(STDOUT). We all know the Hadoop Framework is completely written in java but programs for Hadoop are not necessarily need to code in Java programming language. feature of Hadoop Streaming is available since Hadoop version 0.14.1.

In the above example image, we can see that the flow shown in a dotted block is a basic MapReduce job. In that, we have an Input Reader which is responsible for reading the input data and produces the list of key-value pairs. We can read data in .csv format, in delimiter format, from a database table, image data(.jpg, .png), audio data etc. The only requirement to read all these types of data is that we have to create a particular input format for that data with these input readers. The input reader contains the complete logic about the data it is reading. Suppose we want to read an image then we have to specify the logic in the input reader so that it can read that image data and finally it will generate key-value pairs for that image data.

If we are reading an image data then we can generate key-value pair for each pixel where the key will be the location of the pixel and the value will be its color value from (0-255) for a colored image. Now this list of key-value pairs is fed to the Map phase and Mapper will work on each of these key-value pair of each pixel and generate some intermediate key-value pairs which are then fed to the Reducer after doing shuffling and sorting then the final output produced by the reducer will be written to the HDFS. These are how a simple Map-Reduce job works.

Now let’s see how we can use different languages like Python, C++, Ruby with Hadoop for execution. We can run this arbitrary language by running them as a separate process. For that, we will create our external mapper and run it as an external separate process. These external map processes are not part of the basic MapReduce flow. This external mapper will take input from STDIN and produce output to STDOUT. As the key-value pairs are passed to the internal mapper the internal mapper process will send these key-value pairs to the external mapper where we have written our code in some other language like with python with help of STDIN. Now, these external mappers process these key-value pairs and generate intermediate key-value pairs with help of STDOUT and send it to the internal mappers. 

What is HDFS

Hadoop comes with a distributed file system called HDFS. In HDFS data is distributed over several machines and replicated to ensure their durability to failure and high availability to parallel application.

It is cost effective as it uses commodity hardware. It involves the concept of blocks, data nodes and node name.

Where to use HDFS

  • Very Large Files: Files should be of hundreds of megabytes, gigabytes or more.
  • Streaming Data Access: The time to read whole data set is more important than latency in reading the first. HDFS is built on write-once and read-many-times pattern.
  • Commodity Hardware:It works on low cost hardware.

Where not to use HDFS

  • Low Latency data access: Applications that require very less time to access the first data should not use HDFS as it is giving importance to whole data rather than time to fetch the first record.
  • Lots Of Small Files:The name node contains the metadata of files in memory and if the files are small in size it takes a lot of memory for name node's memory which is not feasible.
  • Multiple Writes:It should not be used when we have to write multiple times.

HDFS Concepts

  1. Blocks: A Block is the minimum amount of data that it can read or write.HDFS blocks are 128 MB by default and this is configurable.Files n HDFS are broken into block-sized chunks,which are stored as independent units.Unlike a file system, if the file is in HDFS is smaller than block size, then it does not occupy full block?s size, i.e. 5 MB of file stored in HDFS of block size 128 MB takes 5MB of space only.The HDFS block size is large just to minimize the cost of seek.
  2. Name Node: HDFS works in master-worker pattern where the name node acts as master.Name Node is controller and manager of HDFS as it knows the status and the metadata of all the files in HDFS; the metadata information being file permission, names and location of each block.The metadata are small, so it is stored in the memory of name node,allowing faster access to data. Moreover the HDFS cluster is accessed by multiple clients concurrently,so all this information is handled bya single machine. The file system operations like opening, closing, renaming etc. are executed by it.
  3. Data Node: They store and retrieve blocks when they are told to; by client or name node. They report back to name node periodically, with list of blocks that they are storing. The data node being a commodity hardware also does the work of block creation, deletion and replication as stated by the name node.


Hadoop Operation

  1. Open cmd in Administrative mode and move to “C:/Hadoop-2.8.0/sbin” and start cluster
Start-all.cmd


  1. Create an input directory in HDFS.
hadoop fs -mkdir /input_dir
  1. Copy the input text file named input_file.txt in the input directory (input_dir)of HDFS.
hadoop fs -put C:/input_file.txt /input_dir
  1. Verify input_file.txt available in HDFS input directory (input_dir).
hadoop fs -ls /input_dir/
  1. Verify content of the copied file.

    hadoop dfs -cat /input_dir/input_file.txt


  1. Run MapReduceClient.jar and also provide input and out directories.
hadoop jar C:/MapReduceClient.jar wordcount /input_dir /output_dir

  1. Verify content for generated output file.
hadoop dfs -cat /output_dir/*

Some Other usefull commands

To leave Safe mode

hadoop dfsadmin –safemode leave

To Delete file from HDFS directory

hadoop fs -rm -r /iutput_dir/input_file.txt

To Delete directory from HDFS directory

hadoop fs -rm -r /iutput_dir



Machine Learning Tutorial

Machine Learning Tutorial

The Machine Learning Tutorial covers both the fundamentals and more complex ideas of machine learning. Students and professionals in the workforce can benefit from our machine learning tutorial.

A rapidly developing field of technology, machine learning allows computers to automatically learn from previous data. For building mathematical models and making predictions based on historical data or information, machine learning employs a variety of algorithms. It is currently being used for a variety of tasks, including speech recognition, email filtering, auto-tagging on Facebook, a recommender system, and image recognition.

You will learn about the many different methods of machine learning, including reinforcement learning, supervised learning, and unsupervised learning, in this machine learning tutorial. Regression and classification models, clustering techniques, hidden Markov models, and various sequential models will all be covered.

What is Machine Learning

In the real world, we are surrounded by humans who can learn everything from their experiences with their learning capability, and we have computers or machines which work on our instructions. But can a machine also learn from experiences or past data like a human does? So here comes the role of Machine Learning.

Introduction to Machine Learning

Introduction to Machine Learning

A subset of artificial intelligence known as machine learning focuses primarily on the creation of algorithms that enable a computer to independently learn from data and previous experiences. Arthur Samuel first used the term "machine learning" in 1959. It could be summarized as follows:

Without being explicitly programmed, machine learning enables a machine to automatically learn from data, improve performance from experiences, and predict things.

Machine learning algorithms create a mathematical model that, without being explicitly programmed, aids in making predictions or decisions with the assistance of sample historical data, or training data. For the purpose of developing predictive models, machine learning brings together statistics and computer science. Algorithms that learn from historical data are either constructed or utilized in machine learning. The performance will rise in proportion to the quantity of information we provide.

A machine can learn if it can gain more data to improve its performance.

How does Machine Learning work

A machine learning system builds prediction models, learns from previous data, and predicts the output of new data whenever it receives it. The amount of data helps to build a better model that accurately predicts the output, which in turn affects the accuracy of the predicted output.

Let's say we have a complex problem in which we need to make predictions. Instead of writing code, we just need to feed the data to generic algorithms, which build the logic based on the data and predict the output. Our perspective on the issue has changed as a result of machine learning. The Machine Learning algorithm's operation is depicted in the following block diagram:

Introduction to Machine Learning

Features of Machine Learning:

  • Machine learning uses data to detect various patterns in a given dataset.
  • It can learn from past data and improve automatically.
  • It is a data-driven technology.
  • Machine learning is much similar to data mining as it also deals with the huge amount of the data.

Need for Machine Learning

The demand for machine learning is steadily rising. Because it is able to perform tasks that are too complex for a person to directly implement, machine learning is required. Humans are constrained by our inability to manually access vast amounts of data; as a result, we require computer systems, which is where machine learning comes in to simplify our lives.

By providing them with a large amount of data and allowing them to automatically explore the data, build models, and predict the required output, we can train machine learning algorithms. The cost function can be used to determine the amount of data and the machine learning algorithm's performance. We can save both time and money by using machine learning.

The significance of AI can be handily perceived by its utilization's cases, Presently, AI is utilized in self-driving vehicles, digital misrepresentation identification, face acknowledgment, and companion idea by Facebook, and so on. Different top organizations, for example, Netflix and Amazon have constructed AI models that are utilizing an immense measure of information to examine the client interest and suggest item likewise.

Following are some key points which show the importance of Machine Learning:

  • Rapid increment in the production of data
  • Solving complex problems, which are difficult for a human
  • Decision making in various sector including finance
  • Finding hidden patterns and extracting useful information from data.

Classification of Machine Learning

At a broad level, machine learning can be classified into three types:

  1. Supervised learning
  2. Unsupervised learning
  3. Reinforcement learning

Introduction to Machine Learning

1) Supervised Learning

In supervised learning, sample labeled data are provided to the machine learning system for training, and the system then predicts the output based on the training data. The system uses labeled data to build a model that understands the datasets and learns about each one. After the training and processing are done, we test the model with sample data to see if it can accurately predict the output.

The mapping of the input data to the output data is the objective of supervised learning. The managed learning depends on oversight, and it is equivalent to when an understudy learns things in the management of the educator. Spam filtering is an example of supervised learning.

Supervised learning can be grouped further in two categories of algorithms:

  • Classification
  • Regression

2) Unsupervised Learning

Unsupervised learning is a learning method in which a machine learns without any supervision.

The training is provided to the machine with the set of data that has not been labeled, classified, or categorized, and the algorithm needs to act on that data without any supervision. The goal of unsupervised learning is to restructure the input data into new features or a group of objects with similar patterns. In unsupervised learning, we don't have a predetermined result. The machine tries to find useful insights from the huge amount of data. It can be further classifieds into two categories of algorithms:

  • Clustering
  • Association

3) Reinforcement Learning

Reinforcement learning is a feedback-based learning method, in which a learning agent gets a reward for each right action and gets a penalty for each wrong action. The agent learns automatically with these feedbacks and improves its performance. In reinforcement learning, the agent interacts with the environment and explores it. The goal of an agent is to get the most reward points, and hence, it improves its performance.

The robotic dog, which automatically learns the movement of his arms, is an example of Reinforcement learning.

Note: We will learn about the above types of machine learning in detail in later chapters.

History of Machine Learning

Before some years (about 40-50 years), machine learning was science fiction, but today it is the part of our daily life. Machine learning is making our day to day life easy from self-driving cars to Amazon virtual assistant "Alexa". However, the idea behind machine learning is so old and has a long history. Below some milestones are given which have occurred in the history of machine learning:

History of Machine Learning

The early history of Machine Learning (Pre-1940):

  • 1834: In 1834, Charles Babbage, the father of the computer, conceived a device that could be programmed with punch cards. However, the machine was never built, but all modern computers rely on its logical structure.
  • 1936: In 1936, Alan Turing gave a theory that how a machine can determine and execute a set of instructions.

The era of stored program computers:

  • 1940: In 1940, the first manually operated computer, "ENIAC" was invented, which was the first electronic general-purpose computer. After that stored program computer such as EDSAC in 1949 and EDVAC in 1951 were invented.
  • 1943: In 1943, a human neural network was modeled with an electrical circuit. In 1950, the scientists started applying their idea to work and analyzed how human neurons might work.

Computer machinery and intelligence:

  • 1950: In 1950, Alan Turing published a seminal paper, "Computer Machinery and Intelligence," on the topic of artificial intelligence. In his paper, he asked, "Can machines think?"

Machine intelligence in Games:

  • 1952: Arthur Samuel, who was the pioneer of machine learning, created a program that helped an IBM computer to play a checkers game. It performed better more it played.
  • 1959: In 1959, the term "Machine Learning" was first coined by Arthur Samuel.

The first "AI" winter:

  • The duration of 1974 to 1980 was the tough time for AI and ML researchers, and this duration was called as AI winter.
  • In this duration, failure of machine translation occurred, and people had reduced their interest from AI, which led to reduced funding by the government to the researches.

Machine Learning from theory to reality

  • 1959: In 1959, the first neural network was applied to a real-world problem to remove echoes over phone lines using an adaptive filter.
  • 1985: In 1985, Terry Sejnowski and Charles Rosenberg invented a neural network NETtalk, which was able to teach itself how to correctly pronounce 20,000 words in one week.
  • 1997: The IBM's Deep blue intelligent computer won the chess game against the chess expert Garry Kasparov, and it became the first computer which had beaten a human chess expert.

Machine Learning at 21st century

2006:

  • Geoffrey Hinton and his group presented the idea of profound getting the hang of utilizing profound conviction organizations.
  • The Elastic Compute Cloud (EC2) was launched by Amazon to provide scalable computing resources that made it easier to create and implement machine learning models.

2007:

  • Participants were tasked with increasing the accuracy of Netflix's recommendation algorithm when the Netflix Prize competition began.
  • Support learning made critical progress when a group of specialists utilized it to prepare a PC to play backgammon at a top-notch level.

2008:

  • Google delivered the Google Forecast Programming interface, a cloud-based help that permitted designers to integrate AI into their applications.
  • Confined Boltzmann Machines (RBMs), a kind of generative brain organization, acquired consideration for their capacity to demonstrate complex information conveyances.

2009:

  • Profound learning gained ground as analysts showed its viability in different errands, including discourse acknowledgment and picture grouping.
  • The expression "Large Information" acquired ubiquity, featuring the difficulties and open doors related with taking care of huge datasets.

2010:

  • The ImageNet Huge Scope Visual Acknowledgment Challenge (ILSVRC) was presented, driving progressions in PC vision, and prompting the advancement of profound convolutional brain organizations (CNNs).

2011:

  • On Jeopardy! IBM's Watson defeated human champions., demonstrating the potential of question-answering systems and natural language processing.

2012:

  • AlexNet, a profound CNN created by Alex Krizhevsky, won the ILSVRC, fundamentally further developing picture order precision and laying out profound advancing as a predominant methodology in PC vision.
  • Google's Cerebrum project, drove by Andrew Ng and Jeff Dignitary, utilized profound figuring out how to prepare a brain organization to perceive felines from unlabeled YouTube recordings.

2013:

  • Ian Goodfellow introduced generative adversarial networks (GANs), which made it possible to create realistic synthetic data.
  • Google later acquired the startup DeepMind Technologies, which focused on deep learning and artificial intelligence.

2014:

  • Facebook presented the DeepFace framework, which accomplished close human precision in facial acknowledgment.
  • AlphaGo, a program created by DeepMind at Google, defeated a world champion Go player and demonstrated the potential of reinforcement learning in challenging games.

2015:

  • Microsoft delivered the Mental Toolbox (previously known as CNTK), an open-source profound learning library.
  • The performance of sequence-to-sequence models in tasks like machine translation was enhanced by the introduction of the idea of attention mechanisms.

2016:

  • The goal of explainable AI, which focuses on making machine learning models easier to understand, received some attention.
  • Google's DeepMind created AlphaGo Zero, which accomplished godlike Go abilities to play without human information, utilizing just support learning.

2017:

  • Move learning acquired noticeable quality, permitting pretrained models to be utilized for different errands with restricted information.
  • Better synthesis and generation of complex data were made possible by the introduction of generative models like variational autoencoders (VAEs) and Wasserstein GANs.
  • These are only a portion of the eminent headways and achievements in AI during the predefined period. The field kept on advancing quickly past 2017, with new leap forwards, strategies, and applications arising.

Machine Learning at present:

The field of machine learning has made significant strides in recent years, and its applications are numerous, including self-driving cars, Amazon Alexa, Catboats, and the recommender system. It incorporates clustering, classification, decision tree, SVM algorithms, and reinforcement learning, as well as unsupervised and supervised learning.

Present day AI models can be utilized for making different expectations, including climate expectation, sickness forecast, financial exchange examination, and so on.

Prerequisites

Before learning machine learning, you must have the basic knowledge of followings so that you can easily understand the concepts of machine learning:

  • Fundamental knowledge of probability and linear algebra.
  • The ability to code in any computer language, especially in Python language.
  • Knowledge of Calculus, especially derivatives of single variable and multivariate functions.

Collaborative Filtering

To address some of the limitations of content-based filtering, collaborative filtering uses similarities between users and items simultaneously to provide recommendations. This allows for serendipitous recommendations; that is, collaborative filtering models can recommend an item to user A based on the interests of a similar user B. Furthermore, the embeddings can be learned automatically, without relying on hand-engineering of features.

Consider a movie recommendation system in which the training data consists of a feedback matrix in which:

  • Each row represents a user.
  • Each column represents an item (a movie).

The feedback about movies falls into one of two categories:

  • Explicit— users specify how much they liked a particular movie by providing a numerical rating.
  • Implicit— if a user watches a movie, the system infers that the user is interested.

To simplify, we will assume that the feedback matrix is binary; that is, a value of 1 indicates interest in the movie.

When a user visits the homepage, the system should recommend movies based on both:

  • similarity to movies the user has liked in the past
  • movies that similar users liked

For the sake of illustration, let's hand-engineer some features for the movies described in the following table:

MovieRatingDescription
The Dark Knight RisesPG-13Batman endeavors to save Gotham City from nuclear annihilation in this sequel to The Dark Knight, set in the DC Comics universe.
Harry Potter and the Sorcerer's StonePGA orphaned boy discovers he is a wizard and enrolls in Hogwarts School of Witchcraft and Wizardry, where he wages his first battle against the evil Lord Voldemort.
ShrekPGA lovable ogre and his donkey sidekick set off on a mission to rescue Princess Fiona, who is emprisoned in her castle by a dragon.
The Triplets of BellevillePG-13When professional cycler Champion is kidnapped during the Tour de France, his grandmother and overweight dog journey overseas to rescue him, with the help of a trio of elderly jazz singers.
MementoRAn amnesiac desperately seeks to solve his wife's murder by tattooing clues onto his body.

Suppose we assign to each movie a scalar in [−1,1] that describes whether the movie is for children (negative values) or adults (positive values). Suppose we also assign a scalar to each user in [−1,1] that describes the user's interest in children's movies (closer to -1) or adult movies (closer to +1). The product of the movie embedding and the user embedding should be higher (closer to 1) for movies that we expect the user to like.

Image showing several movies and users arranged along a one-dimensional embedding space. The position of each movie along this axis describes whether this is a children's movie (left) or an adult movie (right). The position of a user describes interest in children or adult movies.

In the diagram below, each checkmark identifies a movie that a particular user watched. The third and fourth users have preferences that are well explained by this feature—the third user prefers movies for children and the fourth user prefers movies for adults. However, the first and second users' preferences are not well explained by this single feature.

Image of a feedback matrix, where a row corresponds to a user, and a column corresponds to a movie. Each user and each movie is mapped to a one-dimensional embedding (as described in the previous figure), such that the product of the two embeddings approximates the ground truth value in the feedback matrix.

One feature was not enough to explain the preferences of all users. To overcome this problem, let's add a second feature: the degree to which each movie is a blockbuster or an arthouse movie. With a second feature, we can now represent each movie with the following two-dimensional embedding:

Image showing several movies and users arranged on a two-dimensional embedding space. The position of each movie along the horizontal axis describes whether this is a children's movie (left) or an adult movie (right); its position along the vertical axis describes whether this is a blockbuster movie (top) or an arthouse movie (bottom). The position of the users reflect their interests in each category.

We again place our users in the same embedding space to best explain the feedback matrix: for each (user, item) pair, we would like the dot product of the user embedding and the item embedding to be close to 1 when the user watched the movie, and to 0 otherwise.

Image of the same feedback matrix. This time, each user and each movie is mapped to a two-dimensional embedding (as described in the previous figure), such that the dot product of the two embeddings approximates the ground truth value in the feedback matrix.

In this example, we hand-engineered the embeddings. In practice, the embeddings can be learned automatically, which is the power of collaborative filtering models. In the next two sections, we will discuss different models to learn these embeddings, and how to train them.

The collaborative nature of this approach is apparent when the model learns the embeddings. Suppose the embedding vectors for the movies are fixed. Then, the model can learn an embedding vector for the users to best explain their preferences. Consequently, embeddings of users with similar preferences will be close together. Similarly, if the embeddings for the users are fixed, then we can learn movie embeddings to best explain the feedback matrix. As a result, embeddings of movies liked by similar users will be close in the embedding space.

Big Data Analytics - Introduction to R

This section is devoted to introduce the users to the R programming language. R can be downloaded from the cran website. For Windows users, it is useful to install rtools and the rstudio IDE.

The general concept behind R is to serve as an interface to other software developed in compiled languages such as C, C++, and Fortran and to give the user an interactive tool to analyze data.

Navigate to the folder of the book zip file bda/part2/R_introduction and open the R_introduction.Rproj file. This will open an RStudio session. Then open the 01_vectors.R file. Run the script line by line and follow the comments in the code. Another useful option in order to learn is to just type the code, this will help you get used to R syntax. In R comments are written with the # symbol.

In order to display the results of running R code in the book, after code is evaluated, the results R returns are commented. This way, you can copy paste the code in the book and try directly sections of it in R.

# Create a vector of numbers 
numbers = c(1, 2, 3, 4, 5) 
print(numbers) 

# [1] 1 2 3 4 5  
# Create a vector of letters 
ltrs = c('a', 'b', 'c', 'd', 'e') 
# [1] "a" "b" "c" "d" "e"  

# Concatenate both  
mixed_vec = c(numbers, ltrs) 
print(mixed_vec) 
# [1] "1" "2" "3" "4" "5" "a" "b" "c" "d" "e"

Let’s analyze what happened in the previous code. We can see it is possible to create vectors with numbers and with letters. We did not need to tell R what type of data type we wanted beforehand. Finally, we were able to create a vector with both numbers and letters. The vector mixed_vec has coerced the numbers to character, we can see this by visualizing how the values are printed inside quotes.

The following code shows the data type of different vectors as returned by the function class. It is common to use the class function to "interrogate" an object, asking him what his class is.

### Evaluate the data types using class

### One dimensional objects 
# Integer vector 
num = 1:10 
class(num) 
# [1] "integer"  

# Numeric vector, it has a float, 10.5 
num = c(1:10, 10.5) 
class(num) 
# [1] "numeric"  

# Character vector 
ltrs = letters[1:10] 
class(ltrs) 
# [1] "character"  

# Factor vector 
fac = as.factor(ltrs) 
class(fac) 
# [1] "factor"

R supports two-dimensional objects also. In the following code, there are examples of the two most popular data structures used in R: the matrix and data.frame.

# Matrix
M = matrix(1:12, ncol = 4) 
#      [,1] [,2] [,3] [,4] 
# [1,]    1    4    7   10 
# [2,]    2    5    8   11 
# [3,]    3    6    9   12 
lM = matrix(letters[1:12], ncol = 4) 
#     [,1] [,2] [,3] [,4] 
# [1,] "a"  "d"  "g"  "j"  
# [2,] "b"  "e"  "h"  "k"  
# [3,] "c"  "f"  "i"  "l"   

# Coerces the numbers to character 
# cbind concatenates two matrices (or vectors) in one matrix 
cbind(M, lM) 
#     [,1] [,2] [,3] [,4] [,5] [,6] [,7] [,8] 
# [1,] "1"  "4"  "7"  "10" "a"  "d"  "g"  "j"  
# [2,] "2"  "5"  "8"  "11" "b"  "e"  "h"  "k"  
# [3,] "3"  "6"  "9"  "12" "c"  "f"  "i"  "l"   

class(M) 
# [1] "matrix" 
class(lM) 
# [1] "matrix"  

# data.frame 
# One of the main objects of R, handles different data types in the same object.  
# It is possible to have numeric, character and factor vectors in the same data.frame  

df = data.frame(n = 1:5, l = letters[1:5]) 
df 
#   n l 
# 1 1 a 
# 2 2 b 
# 3 3 c 
# 4 4 d 
# 5 5 e 

As demonstrated in the previous example, it is possible to use different data types in the same object. In general, this is how data is presented in databases, APIs part of the data is text or character vectors and other numeric. In is the analyst job to determine which statistical data type to assign and then use the correct R data type for it. In statistics we normally consider variables are of the following types −

  • Numeric
  • Nominal or categorical
  • Ordinal

In R, a vector can be of the following classes −

  • Numeric - Integer
  • Factor
  • Ordered Factor

R provides a data type for each statistical type of variable. The ordered factor is however rarely used, but can be created by the function factor, or ordered.

The following section treats the concept of indexing. This is a quite common operation, and deals with the problem of selecting sections of an object and making transformations to them.

# Let's create a data.frame
df = data.frame(numbers = 1:26, letters) 
head(df) 
#      numbers  letters 
# 1       1       a 
# 2       2       b 
# 3       3       c 
# 4       4       d 
# 5       5       e 
# 6       6       f 

# str gives the structure of a data.frame, it’s a good summary to inspect an object 
str(df) 
#   'data.frame': 26 obs. of  2 variables: 
#   $ numbers: int  1 2 3 4 5 6 7 8 9 10 ... 
#   $ letters: Factor w/ 26 levels "a","b","c","d",..: 1 2 3 4 5 6 7 8 9 10 ...  

# The latter shows the letters character vector was coerced as a factor. 
# This can be explained by the stringsAsFactors = TRUE argumnet in data.frame 
# read ?data.frame for more information  

class(df) 
# [1] "data.frame"  

### Indexing
# Get the first row 
df[1, ] 
#     numbers  letters 
# 1       1       a  

# Used for programming normally - returns the output as a list 
df[1, , drop = TRUE] 
# $numbers 
# [1] 1 
#  
# $letters 
# [1] a 
# Levels: a b c d e f g h i j k l m n o p q r s t u v w x y z  

# Get several rows of the data.frame 
df[5:7, ] 
#      numbers  letters 
# 5       5       e 
# 6       6       f 
# 7       7       g  

### Add one column that mixes the numeric column with the factor column 
df$mixed = paste(df$numbers, df$letters, sep = ’’)  

str(df) 
# 'data.frame': 26 obs. of  3 variables: 
# $ numbers: int  1 2 3 4 5 6 7 8 9 10 ...
# $ letters: Factor w/ 26 levels "a","b","c","d",..: 1 2 3 4 5 6 7 8 9 10 ... 
# $ mixed  : chr  "1a" "2b" "3c" "4d" ...  

### Get columns 
# Get the first column 
df[, 1]  
# It returns a one dimensional vector with that column  

# Get two columns 
df2 = df[, 1:2] 
head(df2)  

#      numbers  letters 
# 1       1       a 
# 2       2       b 
# 3       3       c 
# 4       4       d 
# 5       5       e 
# 6       6       f  

# Get the first and third columns 
df3 = df[, c(1, 3)] 
df3[1:3, ]  

#      numbers  mixed 
# 1       1     1a
# 2       2     2b 
# 3       3     3c  

### Index columns from their names 
names(df) 
# [1] "numbers" "letters" "mixed"   
# This is the best practice in programming, as many times indeces change, but 
variable names don’t 
# We create a variable with the names we want to subset 
keep_vars = c("numbers", "mixed") 
df4 = df[, keep_vars]  

head(df4) 
#      numbers  mixed 
# 1       1     1a 
# 2       2     2b 
# 3       3     3c 
# 4       4     4d 
# 5       5     5e 
# 6       6     6f  

### subset rows and columns 
# Keep the first five rows 
df5 = df[1:5, keep_vars] 
df5 

#      numbers  mixed 
# 1       1     1a 
# 2       2     2b
# 3       3     3c 
# 4       4     4d 
# 5       5     5e  

# subset rows using a logical condition 
df6 = df[df$numbers < 10, keep_vars] 
df6 

#      numbers  mixed 
# 1       1     1a 
# 2       2     2b 
# 3       3     3c 
# 4       4     4d 
# 5       5     5e 
# 6       6     6f 
# 7       7     7g 
# 8       8     8h 
# 9       9     9i 

Various Filesystems in Hadoop


Hadoop is an open-source software framework written in Java along with some shell scripting and C code for performing computation over very large data. Hadoop is utilized for batch/offline processing over the network of so many machines forming a physical cluster. The framework works in such a manner that it is capable enough to provide distributed storage and processing over the same cluster. It is designed to work on cheaper systems commonly known as commodity hardware where each system offers its local storage and computation power.

Hadoop is capable of running various file systems and HDFS is just one single implementation that out of all those file systems. The Hadoop has a variety of file systems that can be implemented concretely. The Java abstract class org.apache.hadoop.fs.FileSystem represents a file system in Hadoop.

Filesystem

URI scheme

Java implementation (all under org.apache.hadoop)

Description

Localfilefs.LocalFileSystemThe Hadoop Local filesystem is used for a locally connected disk with client-side checksumming. The local filesystem uses RawLocalFileSystem with no checksums.
HDFShdfshdfs.DistributedFileSystemHDFS stands for Hadoop Distributed File System and it is drafted for working with MapReduce efficiently. 
HFTPhftphdfs.HftpFileSystem

The HFTP filesystem provides read-only access to HDFS over HTTP. There is no connection of HFTP with FTP. 

This filesystem is commonly used with distcp to share data between HDFS clusters possessing different versions.    

HSFTPhsftphdfs.HsftpFileSystemThe HSFTP filesystem provides read-only access to HDFS over HTTPS. This file system also does not have any connection with FTP.
HARharfs.HarFileSystemThe HAR file system is mainly used to reduce the memory usage of NameNode by registering files in Hadoop HDFS. This file system is layered on some other file system for archiving purposes.
KFS (Cloud-Store)kfsfs.kfs.KosmosFileSystemcloud store or KFS(KosmosFileSystem) is a file system that is written in c++. It is very much similar to a distributed file system like HDFS and GFS(Google File System).
FTPftpfs.ftp.FTPFileSystemThe FTP filesystem is supported by the FTP server.
S3 (native)s3nfs.s3native.NativeS3FileSystemThis file system is backed by AmazonS3.
S3 (block-based)s3fs.s3.S3FileSystemS3 (block-based) file system which is supported by Amazon s3 stores files in blocks(similar to HDFS) just to overcome S3’s file system 5 GB file size limit.  

Hadoop gives numerous interfaces to its various filesystems, and it for the most part utilizes the URI plan to pick the right filesystem example to speak with. You can use any of this filesystem for working with MapReduce while processing very large datasets but distributed file systems with data locality features are preferable like HDFS and KFS(KosmosFileSystem).


Meta Store Services or features

When new data is saved to object storage, we register it into Hive Metastore by calling the metastore API from the code of any data application or orchestration tool. This declarative phase maps a set of objects in the object store to a table exposed by Hive. Part of registration includes specifying the schema of the table held in the file system, with some metadata describing the columns.

Using Hive Metastore in this way provides four main benefits related to:

  1. Virtualization
  2. Discoverability
  3. Schema Evolution
  4. Performance

Let’s discuss these in more detail!

Virtualization

Data analysts using SQL usually aren’t interested in the details of object storage and its access patterns. They’d simply like to have their tables, please!

This dynamic is the driving force that makes Hive Metastore irreplaceable while other Hadoop components were replaced. Every new technology that was introduced made sure to support Hive Metastore to avoid breaking critical analytic workflows dependent upon the table objects defined in Hive.

Discoverability

Hive Metastore naturally becomes a catalog of all the collections held in object storage when exposing new data is accompanied by updating it. If well maintained, this allows for the discovery of data sets available to query.

Additionally, supplemental can be saved in the metastore to provide helpful information about the data like its update frequency, who owns it, etc.

Schema Evolution

One of the challenges of managing data sets over time is their mutability. Records may change over time with respect to the existing columns describing their attributes. Or the set of attributes itself changes over time, resulting in a change to the schema of the table. 

The registration process described above provides a record of the schema for each additional data file that belongs to the table. This means that if the schema has changed at some point in time, it will be recorded within the Hive Metastore. When accessing the data, it can be accessed with the appropriate schema. 

This also provides a good basis to validate a schema if it should not have changed, and alert users to it. Hive holds the information to create such a test.

Performance

Since Hive Metastore maps the table to the underlying object, it allows the representation of a relational database as partitions according to the primary key supported by the object storage. The granularity of the partitions can be set by the user, and if partitions are balanced and their number is reasonable, this mapping allows improvement in query performance.

This is often referred to as “partition pruning”, which allows a query engine to identify data files that can be skipped.


Anatomy of MAP REDUCE

There are five independent entities:

  • The client, which submits the MapReduce job.
  • The YARN resource manager, which coordinates the allocation of
    compute resources on the cluster.
  • The YARN node managers, which launch and monitor the compute
    containers on machines in the cluster.
  • The MapReduce application master, which coordinates the tasks
    running the MapReduce job The application master and the MapReduce tasks run in containers that are scheduled by the resource manager and managed by the node managers.
  • The distributed filesystem, which is used for sharing job files between
    the other entities.

Job Submission :

  • The submit() method on Job creates an internal JobSubmitter
    instance and calls submitJobInternal() on it.
  • Having submitted the job, waitForCompletion polls the job’s
    progress once per second and reports the progress to the console if it
    has changed since the last report.
  • When the job completes successfully, the job counters are displayed
    Otherwise, the error that caused the job to fail is logged to the
    console.

The job submission process implemented by JobSubmitter does the following:

  • Asks the resource manager for a new application ID, used for the
    MapReduce job ID.
  • Checks the output specification of the job For example, if the output
    directory has not been specified or it already exists, the job is not
    submitted and an error is thrown to the MapReduce program.
  • Computes the input splits for the job If the splits cannot be
    computed (because the input paths don’t exist, for example), the job
    is not submitted and an error is thrown to the MapReduce
    program.
  • Copies the resources needed to run the job, including the job
    JAR file, the configuration file, and the computed input splits, to
    the shared filesystem in a directory named after the job ID.
  • Submits the job by calling submitApplication() on the resource
    manager.



Job Initialization :

  • When the resource manager receives a call to its submitApplication() method, it hands off the request to the YARN scheduler.
  • The scheduler allocates a container, and the resource manager then
    launches the application master’s process there, under the node
    manager’s management.
  • The application master for MapReduce jobs is a Java application
    whose main class is MRAppMaster .
  • It initializes the job by creating a number of bookkeeping objects to
    keep track of the job’s progress, as it will receive progress and
    completion reports from the tasks.
  • It retrieves the input splits computed in the client from the shared
    filesystem.
  • It then creates a map task object for each split, as well as a number of
    reduce task objects determined by the mapreduce.job.reduces property (set by the setNumReduceTasks() method on Job).

Task Assignment:

  • If the job does not qualify for running as an uber task, then the
    application master requests containers for all the map and reduce
    tasks in the job from the resource manager .
  • Requests for map tasks are made first and with a higher priority than
    those for reduce tasks, since all the map tasks must complete before
    the sort phase of the reduce can start.
  • Requests for reduce tasks are not made until 5% of map tasks have
    completed.

Task Execution:

  • Once a task has been assigned resources for a container on a
    particular node by the resource manager’s scheduler, the application
    master starts the container by contacting the node manager.
  • The task is executed by a Java application whose main class is
    YarnChild. Before it can run the task, it localizes the resources that
    the task needs, including the job configuration and JAR file, and
    any files from the distributed cache.
  • Finally, it runs the map or reduce task.
  • Streaming runs special map and reduce tasks for the purpose of
    launching the user supplied executable and communicating with it.
  • The Streaming task communicates with the process (which may be
    written in any language) using standard input and output streams.
  • During execution of the task, the Java process passes input key value
    pairs to the external process, which runs it through the user defined
    map or reduce function and passes the output key value pairs back to
    the Java process.
  • From the node manager’s point of view, it is as if the child process
    ran the map or reduce code itself.
  • Progress and status updates :

    • MapReduce jobs are long running batch jobs, taking anything from
      tens of seconds to hours to run.
    • A job and each of its tasks have a status, which includes such things
      as the state of the job or task (e g running, successfully completed,
      failed), the progress of maps and reduces, the values of the job’s
      counters, and a status message or description (which may be set by
      user code).
    • When a task is running, it keeps track of its progress (i e the
      proportion of task is completed).
    • For map tasks, this is the proportion of the input that has been
      processed.
    • For reduce tasks, it’s a little more complex, but the system can still
      estimate the proportion of the reduce input processed.

    It does this by dividing the total progress into three parts,
    corresponding to the three phases of the shuffle.

    • As the map or reduce task runs, the child process communicates
      with its parent application master through the umbilical interface.
    • The task reports its progress and status (including counters) back to
      its application master, which has an aggregate view of the job, every
      three seconds over the umbilical interface.
    • The resource manager web UI displays all the running applications
      with links to the web UIs of their respective application masters,
      each of which displays further details on the MapReduce job,
      including its progress.
    • During the course of the job, the client receives the latest status
      by polling the application master every second (the interval is set
      via mapreduce.client.progressmonitor.pollinterval).
    • Job Completion:

      • When the application master receives a notification that the last
        task for a job is complete, it changes the status for the job to Successful.
      • Then, when the Job polls for status, it learns that the job has
        completed successfully, so it prints a message to tell the user and
        then returns from the waitForCompletion() .
      • Finally, on job completion, the application master and the task
        containers clean up their working state and the OutputCommitter’s
        commitJob () method is called.
      • Job information is archived by the job history server to enable later
        interrogation by users if desired.




  • HIVE

What is HIVE

Hive is a data warehouse system which is used to analyze structured data. It is built on the top of Hadoop. It was developed by Facebook.

Hive provides the functionality of reading, writing, and managing large datasets residing in distributed storage. It runs SQL like queries called HQL (Hive query language) which gets internally converted to MapReduce jobs.

Using Hive, we can skip the requirement of the traditional approach of writing complex MapReduce programs. Hive supports Data Definition Language (DDL), Data Manipulation Language (DML), and User Defined Functions (UDF).

Features of Hive

These are the following features of Hive:

  • Hive is fast and scalable.
  • It provides SQL-like queries (i.e., HQL) that are implicitly transformed to MapReduce or Spark jobs.
  • It is capable of analyzing large datasets stored in HDFS.
  • It allows different storage types such as plain text, RCFile, and HBase.
  • It uses indexing to accelerate queries.
  • It can operate on compressed data stored in the Hadoop ecosystem.
  • It supports user-defined functions (UDFs) where user can provide its functionality.

Limitations of Hive

  • Hive is not capable of handling real-time data.
  • It is not designed for online transaction processing.
  • Hive queries contain high latency.

PIG

Pig Represents Big Data as data flows. Pig is a high-level platform or tool which is used to process the large datasets. It provides a high-level of abstraction for processing over the MapReduce. It provides a high-level scripting language, known as Pig Latin which is used to develop the data analysis codes. First, to process the data which is stored in the HDFS, the programmers will write the scripts using the Pig Latin Language. Internally Pig Engine(a component of Apache Pig) converted all these scripts into a specific map and reduce task. But these are not visible to the programmers in order to provide a high-level of abstraction. Pig Latin and Pig Engine are the two main components of the Apache Pig tool. The result of Pig always stored in the HDFS. 

 

Note: Pig Engine has two type of execution environment i.e. a local execution environment in a single JVM (used when dataset is small in size)and distributed execution environment in a Hadoop Cluster. 

 

Need of Pig: One limitation of MapReduce is that the development cycle is very long. Writing the reducer and mapper, compiling packaging the code, submitting the job and retrieving the output is a time-consuming task. Apache Pig reduces the time of development using the multi-query approach. Also, Pig is beneficial for programmers who are not from Java background. 200 lines of Java code can be written in only 10 lines using the Pig Latin language. Programmers who have SQL knowledge needed less effort to learn Pig Latin. 

·         It uses query approach which results in reducing the length of the code.

·         Pig Latin is SQL like language.

·         It provides many builtIn operators.

·         It provides nested data types (tuples, bags, map).

 

Evolution of Pig: Earlier in 2006, Apache Pig was developed by Yahoo’s researchers. At that time, the main idea to develop Pig was to execute the MapReduce jobs on extremely large datasets. In the year 2007, it moved to Apache Software Foundation(ASF) which makes it an open source project. The first version(0.1) of Pig came in the year 2008. The latest version of Apache Pig is 0.18 which came in the year 2017.

 

Features of Apache Pig: 

·         For performing several operations Apache Pig provides rich sets of operators like the filtering, joining, sorting, aggregation etc.

·         Easy to learn, read and write. Especially for SQL-programmer, Apache Pig is a boon.

·         Apache Pig is extensible so that you can make your own process and  user-defined functions(UDFs) written in python, java or other programming languages .

·         Join operation is easy in Apache Pig.

·         Fewer lines of code.

·         Apache Pig allows splits in the pipeline.

·         By integrating with other components of the Apache Hadoop ecosystem, such as Apache Hive, Apache Spark, and Apache ZooKeeper, Apache Pig enables users to take advantage of these components’ capabilities while transforming data.

·         The data structure is multivalued, nested, and richer.

·         Pig can handle the analysis of both structured and unstructured data.

 

Applications of Apache Pig:  

·         For exploring large datasets Pig Scripting is used.

·         Provides the supports across large data-sets for Ad-hoc queries.

·         In the prototyping of large data-sets processing algorithms.

·         Required to process the time sensitive data loads.

·         For collecting large amounts of datasets in form of search logs and web crawls.

·         Used where the analytical insights are needed using the sampling.

Types of Data Models in Apache Pig: It consist of the 4 types of data models as follows:  

·         Atom: It is a atomic data value which is used to store as a string. The main use of this model is that it can be used as a number and as well as a string.

·         Tuple: It is an ordered set of the fields.

·         Bag: It is a collection of the tuples.

·         Map: It is a set of key/value pairs.

 



BIG SQL

Big SQL enables users to query Hive and HBase data using ANSI compliant SQL. While Hadoop is highly scalable, Big SQL’s advanced cost-based optimizer and Massively Parallel Processing (MPP) architecture as shown in Figure 1, can run queries smarter, not harder, supporting more concurrent users and more complex SQL with less hardware compared to other SQL solutions for Hadoop.

Big SQL is also the ultimate platform for data warehouse offload and consolidation, a key use case for many Hadoop users. This is because Big SQL is the first and only SQL-on-Hadoop solution to understand commonly used SQL syntax from other vendors and products.

 


Need high-performance scans?High-performance insert/updates/deletes? Need machine learning or graph analytics with Spark, with a single security model?

 

Big SQL is a SQL engine for Hadoop that concurrently exploits Hive, HBase and Spark using a single database connection — even a single query

 

The best part of the technology is that all the data belongs to Hadoop. Big SQL tables are Hive tables, Hbase tables, or Spark resilient distributed datasets (RDDs) and integrated with Hive metastore. Save costs by off-loading data to Hadoop and exploit next-generation, fit-for-purpose analytics, with the SQL optimizer for Hadoop.

Following are the features og BigSQL

— It is not for hadoop tables


— Updates work on Big SQL


— It is compatible with Native Tables and HBase Tables — meaning HBase Indexes can be CRUD’ed (loaded too)


— Stored Procedures (SQL) are allowed.


— JDBC/ODBC Driver Support


— ANSI Standardized SQL


— Views supported

 AJIs …It has Big SQL User-Aggregate Join Functions.

 



43 comments:

  1. web designer in Dubai to create amazing websites. Boost your online presence and grow your business with our help. Enhance your brand today. web designer in Dubai

    ReplyDelete
  2. https://clients1.google.com.mx/url?q=https://islamicwalldecors.com%2F
    https://cse.google.ch/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.ch/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.ch/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.fi/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.fi/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://cse.google.fi/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.com.vn/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.com.vn/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://cse.google.com.vn/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.pt/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://clients1.google.pt/url?q=https://islamicwalldecors.com%2F
    https://cse.google.pt/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.com.ua/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://clients1.google.com.ua/url?q=https://islamicwalldecors.com%2F
    https://cse.google.com.ua/url?q=https://islamicwalldecors.com%2F
    https://ipv4.google.com/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://cse.google.ro/url?q=https://islamicwalldecors.com%2F
    https://cse.google.ro/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.ro/url?q=https://islamicwalldecors.com%2F
    https://clients1.google.ro/url?q=https://islamicwalldecors.com/collections/art-prints%2F
    https://cse.google.com.my/url?q=https://islamicwalldecors.com/collections/art-prints/
    https://clients1.google.co.za/url?q=https://islamicwalldecors.com%2F

    ReplyDelete
  3. https://cse.google.be/url?q=https://frenco.ae%2F
    https://clients1.google.be/url?q=https://frenco.ae%2F
    https://clients1.google.be/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.co.th/url?q=https://frenco.ae/web-designer-in-dubai/
    https://cse.google.co.th/url?q=https://frenco.ae%2F
    https://clients1.google.co.th/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.co.th/url?q=https://frenco.ae%2F
    https://clients1.google.com.tr/url?q=https://frenco.ae%2F
    https://clients1.google.com.tr/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.com.tr/url?q=https://frenco.ae%2F
    https://clients1.google.at/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.at/url?q=https://frenco.ae%2F
    https://cse.google.at/url?q=https://frenco.ae%2F
    https://cse.google.cz/url?q=https://frenco.ae%2F
    https://clients1.google.cz/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.cz/url?q=https://frenco.ae%2F
    https://cse.google.se/url?q=https://frenco.ae%2F
    https://clients1.google.se/url?q=https://frenco.ae%2F
    https://clients1.google.se/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.com.mx/url?q=https://frenco.ae%2F
    https://clients1.google.com.mx/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.com.mx/url?q=https://frenco.ae%2F
    https://cse.google.ch/url?q=https://frenco.ae%2F
    https://clients1.google.ch/url?q=https://frenco.ae%2F
    https://clients1.google.ch/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.fi/url?q=https://frenco.ae%2F
    https://clients1.google.fi/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.fi/url?q=https://frenco.ae%2F
    https://clients1.google.com.vn/url?q=https://frenco.ae%2F
    https://clients1.google.com.vn/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.com.vn/url?q=https://frenco.ae%2F
    https://clients1.google.pt/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.pt/url?q=https://frenco.ae%2F
    https://cse.google.pt/url?q=https://frenco.ae%2F
    https://clients1.google.com.ua/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.com.ua/url?q=https%3A%2F%2Ftechsslash.net%2F
    https://cse.google.com.ua/url?q=https://frenco.ae%2F
    https://ipv4.google.com/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.ro/url?q=https://frenco.ae%2F
    https://clients1.google.ro/url?q=https://frenco.ae%2F
    https://clients1.google.ro/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.com.my/url?q=https://frenco.ae/web-designer-in-dubai/
    https://clients1.google.co.za/url?q=https://frenco.ae%2F
    https://clients1.google.co.za/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.co.za/url?q=https://frenco.ae%2F
    https://cse.google.com.sg/url?q=https://frenco.ae%2F
    https://clients1.google.com.sg/url?q=https://frenco.ae%2F
    https://clients1.google.com.sg/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://cse.google.co.kr/url?q=https://frenco.ae%2F
    https://clients1.google.co.kr/url?q=https://frenco.ae/web-designer-in-dubai%2F
    https://clients1.google.co.kr/url?q=https://frenco.ae%2F
    https://cse.google.gr/url?q=https://frenco.ae%2F
    https://cse.google.gr/url?q=https://frenco.ae/web-designer-in-dubai/
    https://clients1.google.gr/url?q=https://frenco.ae/web-designer-in-dubai/

    ReplyDelete
  4. https://clients1.google.ki/url?q=https://frenco.ae/
    https://cse.google.ki/url?q=https://frenco.ae%2F
    https://cse.google.mw/url?q=https://frenco.ae%2F
    https://clients1.google.mw/url?q=https://frenco.ae%2F
    https://clients1.google.mw/url?q=https://frenco.ae%2F
    https://clients1.google.mw/url?q=https://frenco.ae/
    https://clients1.google.ml/url?q=https://frenco.ae%2F
    https://clients1.google.ml/url?q=https://frenco.ae%2F
    https://cse.google.ml/url?q=https://frenco.ae%2F
    https://clients1.google.com.kw/url?q=https://frenco.ae%2F
    https://clients1.google.com.kw/url?q=https://frenco.ae%2F
    https://cse.google.com.kw/url?q=https://frenco.ae%2F
    https://clients1.google.com.lb/url?q=https://frenco.ae%2F
    https://clients1.google.com.lb/url?q=https://frenco.ae%2F
    https://clients1.google.com.lb/url?q=https://frenco.ae/
    https://cse.google.com.lb/url?q=https://frenco.ae%2F
    https://cse.google.is/url?q=https://frenco.ae%2F
    https://cse.google.is/url?q=https://frenco.ae/
    https://clients1.google.is/url?q=https://frenco.ae%2F
    https://clients1.google.is/url?q=https://frenco.ae%2F
    https://clients1.google.is/url?q=https://frenco.ae/
    https://cse.google.co.ls/url?q=https://frenco.ae%2F
    https://clients1.google.co.ls/url?q=https://frenco.ae%2F
    https://cse.google.tt/url?q=https://frenco.ae%2F
    https://clients1.google.tt/url?q=https://frenco.ae%2F
    https://cse.google.ms/url?q=https://frenco.ae%2F
    https://clients1.google.ms/url?q=https://frenco.ae%2F
    https://clients1.google.ms/url?q=https://frenco.ae%2F
    https://clients1.google.ci/url?q=https://frenco.ae%2F
    https://clients1.google.ci/url?q=https://frenco.ae%2F
    https://cse.google.ci/url?q=https://frenco.ae%2F
    https://cse.google.ci/url?q=https://frenco.ae%2F
    https://clients1.google.dm/url?q=https://frenco.ae%2F
    https://clients1.google.dm/url?q=https://frenco.ae%2F
    https://clients1.google.com.sb/url?q=https://frenco.ae%2F
    https://cse.google.com.sb/url?q=https://frenco.ae%2F
    https://clients1.google.dz/url?q=https://frenco.ae%2F
    https://clients1.google.dz/url?q=https://frenco.ae%2F
    https://cse.google.dz/url?q=https://frenco.ae%2F
    https://cse.google.co.vi/url?q=https://frenco.ae%2F
    https://clients1.google.co.vi/url?q=https://frenco.ae%2F
    https://clients1.google.co.vi/url?q=https://frenco.ae%2F

    ReplyDelete

02

Capstone resource hub

Codingacharya

Capstone Learning Resources, Notes & Project Hub

TCS NQT Questions
Read Notes
Machine Learning – ACE Theory
Read Notes
Machine Learning PPT
Read Notes
MachienLearning LAB
Read Notes
CSPT LAB programs
Read Notes
Time table and CSPT syllabus
Read Notes
Appreciations
Read Notes
ISTE life memberships
Read Notes
Artificial Intelligence & Analytics
Read Notes
Fullstack Web Dev
Read Notes
MERN Web Dev
Read Notes
Course Structure
Read Notes
Cloud Computing
Read Notes
90 Days ML Challenge
Read Notes
Advanced Analytics & Viz
Read Notes
Advanced Machine Learning
Read Notes
React JS
Read Notes
ML Chaitanya
Read Notes
Important Links
Read Notes
CSS Effects
Read Notes
RESUME
Read Notes
Bootstrap CSS
Read Notes
MongoDB
Read Notes
OWN Python Package
Read Notes
HTML Course
Read Notes
HTML Projects
Read Notes
GitHub Projects
Read Notes
Angular JS
Read Notes
Journals
Read Notes
NLP Notes
Read Notes
Videos
Read Notes
Data Analytics & Viz
Read Notes
Cloud Computing (Archive)
Read Notes
Open CV
Read Notes
jQuery
Read Notes
React JS (Archive)
Read Notes
Node JS
Read Notes
DAV Theory
Read Notes
DAV Lab
Read Notes
Big Data Notes
Read Notes
R-Programming
Read Notes
HADOOP Lab
Read Notes
GATE DA
Read Notes
JAVA Lab
Read Notes
Computer Networks
Read Notes
03

Live projects & profiles