Spark-1.0.0 SQL使用简介

Spark-1.0.0 SQL使用简介_第1张图片

    • 上传文件到HDFS
    • 启动 sql

1.上传文件到HDFS

http://blog.csdn.net/zhaolei5911/article/details/64514726

2.启动 sql

spark1.0.0 中 sql 启动是直接在 spark-shell 启动后 启动

val sqlContext = new org.apache.spark.sql.SQLContext(sc)
import sqlContext._

// Define the schema using a case class.
// Note: Case classes in Scala 2.10 can support only up to 22 fields. To work around this limit, 
// you can use custom classes that implement the Product interface.
case class Person(name: String, age: Int)

// Create an RDD of Person objects and register it as a table.
val people = sc.textFile("hdfs://hadoopmaster:8020/data/wordcount/people.txt").map(_.split(",")).map(p => Person(p(0), p(1).trim.toInt))
people.registerAsTable("people")

// SQL statements can be run by using the sql methods provided by sqlContext.
val teenagers = sql("SELECT name FROM people WHERE age >= 13 AND age <= 19")

// The results of SQL queries are SchemaRDDs and support all the normal RDD operations.
// The columns of a row in the result can be accessed by ordinal.
teenagers.map(t => "Name: " + t(0)).collect().foreach(println)

people是txt文件,内容形式如下:
aaa 19
bbb 29

运行结果如下:

Spark-1.0.0 SQL使用简介_第2张图片

你可能感兴趣的:(spark,sql,spark-sql)