Hadoop Installation
- https://www.virtualbox.org/wiki/Downloads
- https://www.youtube.com/redirect?event=video_description&redir_token=QUFFLUhqbjVXblY2dmpxNDBkX2wxMTZ1aDZiVndHTjF2Z3xBQ3Jtc0tsNm5sT01GS3htRlc1LXlra3NmSGdfNktLQ2xrMXc4Z3l0cUNwSU95MmMweXBwYWdrOHFna0JaVHJmbGJoMU9TZjhIVFY2YVloaGVDYmc3YXktcFRBRFRhVG40clhocURyN0lubUZ5MzhVNl9NYkxuTQ&q=https%3A%2F%2Fdownloads.cloudera.com%2Fdemo_vm%2Fvirtualbox%2Fcloudera-quickstart-vm-5.13.0-0-virtualbox.zip&v=8YFZrTQW97A
- https://www.7-zip.org/download.html
After Installing applications open terminal then enter these commands in terminal
$ hostname
$ hdfs dfs -ls /
$ sudo /home/cloudera/cloudera-manager --express --force
copy this link address from terminal
http://quickstart.cloudera:7180
username:cloudera
password:cloudera
Open terminal enter these commands as it is
program 3 (HIVE tool)
open terminal
enter the following commands to create a database in HIVE
$ hive
hive> show databases;
hive> create database geeksportal;
hive> create table geeksportal.geekdata(id int, name string);
Program 4
Open Terminal
Enter these commands
$ hive
hive >
CREATE TABLE IF NOT EXISTS student( Student_Name STRING, Student_Rollno INT, Student_Marks FLOAT) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',';
INSERT INTO TABLE student VALUES ('Dikshant',1,'95'),('Akshat', 2 , '96'),('Dhruv',3,'90');SELECT * FROM student;
Program 5
Q: How to upload a CSV file in to hive(using hue)
open hue
enter username and passwords (cloudera, cloudera)
click on file browser
click on upload and select files
click on data browsers from menu bar then choose meta store tables
then click on create a new table from a file
enter table name
choose csv file from input file
select delimiter like comma
click on use first row as a column name
then click on create table
Program 6
Find the salary of an employee > 30000 using HIVE.
step 1: create database sample_11;
step 2: CREATE TABLE IF NOT EXISTS company(code STRING, description STRING, total_emp STRING, salary STRING) ROW FORMAT DELIMITED FIELDS TERMINATED BY ',';
step 3: INSERT INTO TABLE company VALUES('001','SATISH','100','30500');
step 4: SELECT * FROM company;
step 5: SELECT company.description, company.salary FROM company WHERE(company.salary > 30000) ORDER BY company.salary DESC LIMIT 1000;
Program 7
A simple program using SQOOP
$ mysql -u root -p;
mysql> show databases;
mysql> use retails_db;
mysql> show tables;
mysql> select * from customers;
mysql> describe customers;
Now open new terminal
$ hive
hive> show databases;
hive> create database export_db;
hive> use export_db;
hive> create table customers(customer_id varchar(45), customer_fname varchar(45), customer_lname varchar(45), customer_email varchar(45), customer_password varchar(45), customer_street(255), customer_city(45), customer_state(45), customer_zipcode(45));
hive> show tables;
hive> select count(*) from customers;
Now open New Terminal
$ sqoop import --connect jdbc:mysql://localhost/retail_db --username=root --password=cloudera --table=customers --hive-home=/user/hive/warehouse --hive-import --hive-overwrite --hive-table=export_db.customers
Now enter this command in HIVE
hive> select count(*) from customers limit 10;
https://techvidvan.com/tutorials/apache-pig-operators/ for inner joins and operators.
Hadoop Operation
- Open cmd in Administrative mode and move to “C:/Hadoop-2.8.0/sbin” and start cluster
Start-all.cmd- Create an input directory in HDFS.
hadoop fs -mkdir /input_dir
- Copy the input text file named input_file.txt in the input directory (input_dir)of HDFS.
hadoop fs -put C:/input_file.txt /input_dir
- Verify input_file.txt available in HDFS input directory (input_dir).
hadoop fs -ls /input_dir/- Verify content of the copied file.
- Run MapReduceClient.jar and also provide input and out directories.
hadoop jar C:/MapReduceClient.jar wordcount /input_dir /output_dir- Verify content for generated output file.
hadoop dfs -cat /output_dir/*Some Other usefull commands
To leave Safe mode
hadoop dfsadmin –safemode leave
To Delete file from HDFS directory
hadoop fs -rm -r /iutput_dir/input_file.txt
To Delete directory from HDFS directory
hadoop fs -rm -r /iutput_dir
No comments:
Post a Comment