Hadoop Tool,ToolRunner原理分析

wbj0110

浏览: 1549685 次
性别:
来自: 上海

最近访客更多访客>>

一往无前bhz

ninja2006

loginboot

u012363178

博主相关

博客

微博

相册

留言

关于我

文章分类

社区版块

存档分类

博客分类：

Hadoop

Hadoop

先看Configurable 接口：

public interface Configurable {
void setConf(Configuration conf);
  Configuration getConf();
}

Configurable接口只定义了两个方法：setConf与 getConf。
Configured类实现了Configurable接口：

public class Configured implements Configurable {
  private Configuration conf;
    public Configured() {
    this(null);
  }

  public Configured(Configuration conf) {
    setConf(conf);
  }
 
  public void setConf(Configuration conf) {
    this.conf = conf;
  }
 public Configuration getConf() {
    return conf;
  }
}

Tool接口继承了Configurable接口，只有一个run()方法。(接口继承接口)

public interface Tool extends Configurable {
  int run(String [] args) throws Exception;
}

继承关系如下：

再看ToolRunner类的一部分：

public class ToolRunner {
  public static int run(Configuration conf, Tool tool, String[] args)
  throws Exception{
    if(conf == null) {
     conf = new Configuration();
    }

    GenericOptionsParser parser = new GenericOptionsParser(conf, args);
    //set the configuration back, so that Tool can configure itself
    tool.setConf(conf);
    //get the args w/o generic hadoop args
    String[] toolArgs = parser.getRemainingArgs();
    return tool.run(toolArgs);

  }
}

从ToolRunner的静态方法run()可以看到，其通过GenericOptionsParser 来读取传递给run的job的conf和命令行参数args，处理hadoop的通用命令行参数，然后将剩下的job自己定义的参数(toolArgs = parser.getRemainingArgs();)交给tool来处理,再由tool来运行自己的run方法。

通用命令行参数指的是对任意的一个job都可以添加的，如：

-conf < configuration file >     specify a configuration file
-D < property=value >            use value for given property
-fs < local|namenode:port >      specify a namenode
-jt < local|jobtracker:port >    specify a job tracker
-files < comma separated list of files >    specify comma separated files to be copied to the map reduce cluster
-libjars < comma separated list of jars >   specify comma separated jar files to include in the classpath.
-archives < comma separated list of archives >    specify comma separated archives to be unarchived on the compute machines.

一个典型的实现Tool的程序：

/**

MyApp 需要从命令行读取参数，用户输入命令如，

$bin/hadoop jar MyApp.jar -archives test.tgz  arg1 arg2

-archives 为hadoop通用参数，arg1 ,arg2为job的参数

*/

public class MyApp extends Configured implements Tool {

//implemet Tool’s run

    public int run(String[] args) throws Exception {

        Configuration conf = getConf();

// Create a JobConf using the processed conf

        JobConf job = new JobConf(conf, MyApp.class);

// Process custom command-line options

        Path in = new Path(args[1]);

        Path out = new Path(args[2]);

// Specify various job-specific parameters

        job.setJobName("my-app");

        job.setInputPath(in);

        job.setOutputPath(out);

        job.setMapperClass(MyApp.MyMapper.class);

        job.setReducerClass(MyApp.MyReducer.class);

        JobClient.runJob(job);

    }

    public static void main(String[] args) throws Exception {

// args由ToolRunner来处理

        int res = ToolRunner.run(new Configuration(), new MyApp(), args);

        System.exit(res);

    }

}

http://hnote.org/big-data/hadoop/hadoop-tool-toolrunner

分享到：

Mahout之SparseVectorsFromSequenceFiles源 ... | Hbase shell 常用命令

2014-06-19 09:18
浏览 892
评论(0)
分类:编程语言
查看更多

发表评论

您还没有登录,请您登录后再发表评论

最近访客更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

Hadoop Tool,ToolRunner原理分析

评论

发表评论

相关推荐

最近访客 更多访客>>

博主相关

文章分类

社区版块

存档分类

最新评论

Hadoop Tool,ToolRunner原理分析

评论

发表评论

相关推荐

Hadoop DistributedCache使用及原理

HBase高性能复杂条件查询引擎

HADOOP基本操作命令

在线分析查询系统mdrill

Hadoop实现AbstractJob简化Job设置

让你彻底明白hive数据存储各种模式

YARN 各种RPC通信协议及它们的作用介绍

YARN工作流程

HADOOP工作流调度系统OOZIE

Hadoop 中利用 mapreduce 读写 mysql 数据

hadoop编程：解决eclipse能运行，打包放到集群上ClassNotFoundException:经验总结

分别使用Hadoop MapReduce、hive统计手机流量

eclipse中开发Hadoop2.x的Map/Reduce项目汇总

Cloudera Impala: Real-Time Queries in Apache Hadoop, For Real

Eclipse调用hadoop2运行MR程序

Mahout for hadoop 2

hadoop2.2+mahout0.9实战

STS或eclipse安装SVN插件

大数据入门：各种大数据技术介绍

hadoop开发方式总结及操作指导

最近访客更多访客>>