继承关系:
1. java.util
Interface Map.Entry<K,V>
description:
public static interface Map.Entry<K,V>
methods:
Modifier and Type | Method and Description |
---|---|
boolean |
equals(Object o)
Compares the specified object with this entry for equality.
|
K |
getKey()
Returns the key corresponding to this entry.
|
V |
getValue()
Returns the value corresponding to this entry.
|
int |
hashCode()
Returns the hash code value for this map entry.
|
V |
setValue(V value)
Replaces the value corresponding to this entry with the specified value (optional operation).
|
2.java.lang.Object
|__ org.apache.hadoop.conf.Configuration
constructor:
public class Configuration
extends Objectimplements Iterable<Map.Entry<String,String>>, Writable
3.org.apache.hadoop.util
Class ToolRunner
java.lang.Object
|__ org.apache.hadoop.util.ToolRunner
description:
public class ToolRunner
extends Object
ToolRunner
can be used to run classes implementing Tool
interface. It works in conjunction with GenericOptionsParser
to parse the generic hadoop command line arguments and modifies the Configuration
of the Tool
. The application-specific options are passed along without being modified.
methods:
static int |
run(Configuration conf, Tool tool, String[] args) Runs the given Tool by Tool.run(String[]) , after parsing with the given generic arguments. |
static int |
run(Tool tool, String[] args) Runs the Tool with its Configuration . |
4.org.apache.hadoop.util
Interface Tool
description:
public interface Tool
extends Configurable
methods:
int |
run(String[] args) Execute the command with the given arguments. |
5.org.apache.hadoop.conf
Interface Configurable
constructor:
public interface Configurable
methods:
Configuration |
getConf() Return the configuration used by this object. |
void |
setConf(Configuration conf) Set the configuration to be used by this object. |
6.
java.lang.Object |__ org.apache.hadoop.conf.Configured
description:
public class Configured
extends Object
implements Configurable
constructor:
Configured() Construct a Configured. |
Configured(Configuration conf) Construct a Configured |
methods:
Configuration |
getConf() Return the configuration used by this object. |
void |
setConf(Configuration conf) Set the configuration to be used by this object. |
Code1 (Configuration里添加的resource是String类型):
1 import java.util.Map.Entry; 2 3 import org.apache.hadoop.conf.Configuration; 4 import org.apache.hadoop.conf.Configured; 5 import org.apache.hadoop.util.ToolRunner; 6 import org.apache.hadoop.util.Tool; 7 import org.apache.hadoop.fs.Path; 8 9 public class ConfigurationPrinter extends Configured implements Tool { 10 static { 11 Configuration.addDefaultResource("config.xml"); 12 } 13 14 @Override 15 public int run(String[] args) throws Exception { 16 Configuration conf = getConf(); 17 for (Entry<String, String> hash: conf) { 18 System.out.printf("%s=%s\n", hash.getKey(), hash.getValue()); 19 } 20 return 0; 21 } 22 23 public static void main(String[] args) throws Exception { 24 int exitCode = ToolRunner.run(new ConfigurationPrinter(), args); 25 System.exit(exitCode); 26 } 27 }
注:Configuration class提供只一种静态方法:addDefaultresource(String name), 如上述代码, 添加Resource "config.xml"为String类型时,hadoop将从classpath里查找此文件;若Resource 为Path()类型时,hadoop将从local filesystem里查找此文件: Configuration conf = new Configuration(); conf.addResource(new Path("config.xml"));
code1的执行步骤:
#将自定义的config文件config.xml放在hadoop的$HADOOP_CONF_DIR里 mv config.xml $HADOOP_HOME/etc/hadoop/
#假如我们添加的resource如下:
1 <!--cat $HADOOP_HOME/etc/hadoop/config.xml--> 2 <configuration> 3 <property> 4 <name>color</name> 5 <value>yellow</value> 6 </property> 7 8 <property> 9 <name>size</name> 10 <value>10</value> 11 </property> 12 13 <property> 14 <name>weight</name> 15 <value>heavy</value> 16 <final>true</final> 17 </property> 18 </configuration>
执行代码:
mkdir class source $HADOOP_HOME/libexec/hadoop-config.sh javac -d class ConfigurationPrinter.java jar -cvf ConfigurationPrinter.jar -C class ./ export HADOOP_CLASSPATH=ConfigurationPrinter.jar:$CLASSPATH #下面查找刚才添加的resource是否被读入 #我们在config.xml里添加了一项 <name>color</name>,执行 yarn ConfigurationPrinter|grep "color" color=yellow #可见代码是正确的
或者在commandline里指定HADOOP_CONF_DIR, 比如执行:
yarn ConfigurationPrinter --conf config.xml | grep color
color=yellow
也是可以的!
Code2 (Configuration里添加的resource是Path类型):
1 import java.util.Map.Entry; 2 3 import org.apache.hadoop.conf.Configuration; 4 import org.apache.hadoop.conf.Configured; 5 import org.apache.hadoop.util.ToolRunner; 6 import org.apache.hadoop.util.Tool; 7 import org.apache.hadoop.fs.Path; 8 9 public class ConfigurationPrinter extends Configured implements Tool { 10 @Override 11 public int run(String[] args) throws Exception { 12 Configuration conf = new Configuration(); 13 conf.addResource(new Path("config.xml")); 14 for (Entry<String, String> hash: conf) { 15 System.out.printf("%s=%s\n", hash.getKey(), hash.getValue()); 16 } 17 return 0; 18 } 19 20 public static void main(String[] args) throws Exception { 21 int exitCode = ToolRunner.run(new ConfigurationPrinter(), args); 22 System.exit(exitCode); 23 } 24 }
此时添加的resource类型是Path()类型,故hadoop将从local filesystem里查找config.xml, 不需要将config.xml放在conf/下面,只要在代码中指定config.xml在本地文件系统中的路径即可(new Path("../others/config.xml"))
运行步骤:
mkdir class source $HADOOP_HOME/libexec/hadoop-config.sh javac -d class ConfigurationPrinter.java jar -cvf ConfigurationPrinter.jar -C class ./ export HADOOP_CLASSPATH=ConfigurationPrinter.jar:$CLASSPATH #下面查找刚才添加的resource是否被读入 #我们在config.xml里添加了一项 <name>color</name>,执行 yarn ConfigurationPrinter|grep "color" color=yellow #可见代码是正确的
备注:ConfigurationParser支持set individual properties:
Generic Options The supported generic options are: -conf <configuration file> specify a configuration file -D <property=value> use value for given property -fs <local|namenode:port> specify a namenode -jt <local|jobtracker:port> specify a job tracker -files <comma separated list of files> specify comma separated files to be copied to the map reduce cluster -libjars <comma separated list of jars> specify comma separated jar files to include in the classpath. -archives <comma separated list of archives> specify comma separated archives to be unarchived on the compute machines.
可以尝试:
yarn ConfigurationPrinter -d fuck=Japan | grep fuck #输出为: fuck=Japan
再次提醒:
ToolRunner
can be used to run classes implementing Tool
interface. It works in conjunction with GenericOptionsParser
to parse the generic hadoop command line arguments and modifies the Configuration
of the Tool
. The application-specific options are passed along without being modified.
ToolRunner和GenericOptionsParser共同来(解析|修改) generic hadoop command line arguments (什么是generic hadoop command line arguments? 比如:yarn command [genericOptions] [commandOptions]