flink之max与maxBy的区别

package com.sandra.day03;

import com.atguigu.bean.WaterSensor;
import org.apache.flink.api.common.functions.MapFunction;
import org.apache.flink.api.java.tuple.Tuple;
import org.apache.flink.streaming.api.datastream.KeyedStream;
import org.apache.flink.streaming.api.datastream.SingleOutputStreamOperator;
import org.apache.flink.streaming.api.environment.StreamExecutionEnvironment;

public class Flink07_Transform_MaxAndMaxBy {
    public static void main(String[] args) throws Exception {
        //1.获取流的执行环境
        StreamExecutionEnvironment env = StreamExecutionEnvironment.getExecutionEnvironment();
        env.setParallelism(1);

        //2.从端口读取数据并转为JavaBean
        KeyedStream keyedStream = env.socketTextStream("hadoop102", 9999)
                .map(new MapFunction() {
                    @Override
                    public WaterSensor map(String value) throws Exception {
                        String[] split = value.split(",");
                        return new WaterSensor(split[0], Long.parseLong(split[1]), Integer.parseInt(split[2]));
                    }
                })
                //将相同id的watersensor聚合到一起
                .keyBy("id");
        //todo 3.使用简单聚合滚动算子
        //max
//        SingleOutputStreamOperator result = keyedStream.max("vc");
        //maxBy
//        SingleOutputStreamOperator result = keyedStream.maxBy("vc");//默认是true
        SingleOutputStreamOperator result = keyedStream.maxBy("vc", false);

        result.print("maxBy");

        env.execute();

    }
}

//聚合算子需要做keyBy变成keyedStream之后才能调用

flink之max与maxBy的区别_第1张图片

flink之max与maxBy的区别_第2张图片

总结:

max和maxBy区别:max中非比较字段取值取得是第一条数据的值,maxBy取得时最大值的非比较字段的值。
maxBy-true  当两条数据都是最大值的时候,非比较字段取比较早的那个值 
maxBy-false 当两条数据都是最大值的时候,非比较字段取最新的那个值

聚合算子的作用范围是分组内,是对同一分组的数据做聚合

聚合算子特征来一条聚合一条

你可能感兴趣的:(flink,flink,大数据)