俗话说:磨刀不误砍柴工。。上两篇中,我们介绍完了CDH环境的基本搭建。在这篇中,我们讲述对hive的一个优化措施之一:执行引擎tez。在HDP中hive的执行引擎默认是tez,而在CDH中默认则是mr,大家都知道mr效率是比较低的,tez则相对比较高。
tez官网地址为:http://tez.apache.org/
tez是在hadoop2才引入的,如图所示:
图片来源:https://zh.hortonworks.com/apache/tez/
1.1前置环境
1)安装JDK
2)安装Maven
下载安装包:apache-maven-3.5.4-bin.tar.gz
解压:
[root@cm softs]#
tar -zxvf apache-maven-3.5.4-bin.tar.gz -C /usr/local/software/maven
配置:
[root@cm ~]# vim /etc/profile
export MAVEN_HOME=/usr/local/software/maven/apache-maven-3.5.4
export PATH=${MAVEN_HOME}/bin:$PATH
[root@cm ~]# source /etc/profile
3)安装protobuf-2.5.0
解压:tar -zxvf protobuf-2.5.0.tar.gz
编译:
[root@cm protobuf-2.5.0]# cd protobuf-2.5.0
[root@cm protobuf-2.5.0]# ./configure
[root@cm protobuf-2.5.0]# make
[root@cm protobuf-2.5.0]# make install
[root@cm protobuf-2.5.0]# protoc --version #验证是否安装成功
1.2下载并解压tez
(1)下载地址:http://tez.apache.org/releases/
http://tez.apache.org/releases/apache-tez-0-9-2.html
(2)解压命令:tar -zxvf apache-tez-0.9.2-src.tar.gz
1.3修改
建议在win10下通过编辑工具修改,因为文件比较大,比较难找到需要修改的地方
(1)修改pom.xml
第一处:修改为我们cdh所用版本
第二三处:添加Cloudera的Maven仓库地址【因为Hadoop环境版本为CDH版本】
第四处:注释掉tez-ext-service-tests、tez-ui这两个模块
${distMgmtSnapshotsId}
${distMgmtSnapshotsName}
${distMgmtSnapshotsUrl}
cloudera
https://repository.cloudera.com/artifactory/cloudera-repos/
Cloudera Repositories
false
maven2-repository.atlassian
Atlassian Maven Repository
https://maven.atlassian.com/repository/public
default
${distMgmtSnapshotsId}
${distMgmtSnapshotsName}
${distMgmtSnapshotsUrl}
default
cloudera
Cloudera Repositories
https://repository.cloudera.com/artifactory/cloudera-repos/
(2)修改java类
进入到:cd tez-mapreduce/src/main/java/org/apache/tez/mapreduce/hadoop/mapreduce
修改:JobContextImpl.java,在文件最后新增:
/**
* Get the boolean value for the property that specifies which classpath
* takes precedence when tasks are launched. True - user's classes takes
* precedence. False - system's classes takes precedence.
* @return true if user's classes should take precedence
*/
@Override
public boolean userClassesTakesPrecedence() {
return getJobConf().getBoolean(MRJobConfig.MAPREDUCE_JOB_USER_CLASSPATH_FIRST, false);
}
1.4编译
mvn clean package -DskipTests=true -Dmaven.javadoc.skip=true
1.5和CDH整合
1)进入apache-tez-0.9.2-src/tez-dist/target
将tez-0.9.2.tar.gz压缩包上传到HDFS上:/engine/tez-0.9.2
2)创建tez目录
在/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib 下创建tez目录
进入到tez目录下,创建conf目录
mkdir conf
将1)中的tez-0.9.2-minimal文件夹下的jar及lib下的jar拷贝到tez中
在/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/tez 中如下目录:
其中conf文件夹下,放置tez-site.xml
tez.lib.uris
${fs.defaultFS}/engine/tez-0.9.2/tez-0.9.2.tar.gz
tez.use.cluster.hadoop-libs
false
避免kryo的错误:
java.lang.NoClassDefFoundError: com/esotericsoftware/shaded/org/objenesis/strategy/InstantiatorStrategy
删除或者重命名hive/auxlib下的hive-exec-1.1.0-cdh5.16.1-core.jar和hive-exec-core.jar 否则会有kryo错误
我这里的CDH安装目录为:
/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/hive/auxlib
将两个jar包删除或重名下即可
[root@master1 ~]# cd /opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/hive/auxlib
[root@master1 auxlib]# mv hive-exec-1.1.0-cdh5.16.1-core.jar hive-exec-1.1.0-cdh5.16.1-core.jar.bak
[root@master1 auxlib]# mv hive-exec-core.jar hive-exec-core.jar.bak
3)在CDH的hive中配置
HADOOP_CLASSPATH=/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/tez/conf:/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/tez/*:/opt/cloudera/parcels/CDH-5.16.1-1.cdh5.16.1.p0.3/lib/tez/lib/*
保存配置,重启Hive。。
4)验证