Flume采集目录及文件到HDFS案例

程序员文章站 2022-03-18 21:52:04

采集目录到HDFS 使用flume采集目录需要启动hdfs集群 spooldir source 监控指定目录如果目录下有新文件产生就采集走注意！！！此组件监控的目录不能有同名的文件产生一旦有重名文件：报错罢工注意！！！此组件监控的目录不能有同名的文件产生一旦有重名文件：报错罢工 ......

采集目录到HDFS

　　使用flume采集目录需要启动hdfs集群

vi spool-hdfs.conf

# Name the components on this agent
a1.sources = r1
a1.sinks = k1
a1.channels = c1

# Describe/configure the source
##注意：不能往监控目中重复丢同名文件
a1.sources.r1.type = spooldir
a1.sources.r1.spoolDir = /root/logs2
a1.sources.r1.fileHeader = true

# Describe the sink
a1.sinks.k1.type = hdfs
a1.sinks.k1.channel = c1
a1.sinks.k1.hdfs.path = /flume/events/%y-%m-%d/%H%M/
a1.sinks.k1.hdfs.filePrefix = events-
#控制文件夹的滚动频率
a1.sinks.k1.hdfs.round = true
a1.sinks.k1.hdfs.roundValue = 10
a1.sinks.k1.hdfs.roundUnit = minute
#控制文件的滚动频率
a1.sinks.k1.hdfs.rollInterval = 3  #时间维度
a1.sinks.k1.hdfs.rollSize = 20　　#文件大小维度
a1.sinks.k1.hdfs.rollCount = 5　　#event数量维度
a1.sinks.k1.hdfs.batchSize = 1
a1.sinks.k1.hdfs.useLocalTimeStamp = true
#生成的文件类型，默认是Sequencefile，可用DataStream，则为普通文本
a1.sinks.k1.hdfs.fileType = DataStream

# Use a channel which buffers events in memory
a1.channels.c1.type = memory
a1.channels.c1.capacity = 1000
a1.channels.c1.transactionCapacity = 100

# Bind the source and sink to the channel
a1.sources.r1.channels = c1
a1.sinks.k1.channel = c1

mkdir /root/logs2

　　　　spooldir source 监控指定目录如果目录下有新文件产生就采集走

- 注意！！！此组件监控的目录不能有同名的文件产生一旦有重名文件：报错罢工

　　启动命令：

bin/flume-ng agent -c ./conf -f ./conf/spool-hdfs.conf -n a1 -Dflume.root.logger=INFO,console

采集文件到HDFS

vi tail-hdfs.conf

# Name the components on this agent
a1.sources = r1
a1.sinks = k1
a1.channels = c1

# Describe/configure the source
a1.sources.r1.type = exec
a1.sources.r1.command = tail -F /root/logs/test.log
a1.sources.r1.channels = c1

# Describe the sink
a1.sinks.k1.type = hdfs
a1.sinks.k1.channel = c1
a1.sinks.k1.hdfs.path = /flume/tailout/%y-%m-%d/%H-%M/
a1.sinks.k1.hdfs.filePrefix = events-
a1.sinks.k1.hdfs.round = true
a1.sinks.k1.hdfs.roundValue = 10
a1.sinks.k1.hdfs.roundUnit = minute
a1.sinks.k1.hdfs.rollInterval = 3
a1.sinks.k1.hdfs.rollSize = 20
a1.sinks.k1.hdfs.rollCount = 5
a1.sinks.k1.hdfs.batchSize = 1
a1.sinks.k1.hdfs.useLocalTimeStamp = true
#生成的文件类型，默认是Sequencefile，可用DataStream，则为普通文本
a1.sinks.k1.hdfs.fileType = DataStream



# Use a channel which buffers events in memory
a1.channels.c1.type = memory
a1.channels.c1.capacity = 1000
a1.channels.c1.transactionCapacity = 100

# Bind the source and sink to the channel
a1.sources.r1.channels = c1
a1.sinks.k1.channel = c1

mkdir /root/logs

启动命令

bin/flume-ng agent -c conf -f conf/tail-hdfs.conf -n a1

exec source 可以执行一个shell命令（tail -F sx.log）实时采集文件数据变化

模拟数据生成的脚步：

while true;do date >> /root/logs/test.log;sleep 0.5;done

或

    #!/bin/bash

while true

do

  date >> /root/logs/test.log

  sleep 1

done

上一篇： MapReduce序列化及分区的java代码示例

下一篇： mysql-foreach语句怎么就插入第一个数据

Flume采集目录及文件到HDFS案例

采集目录到HDFS

采集文件到HDFS

Flume采集目录及文件到HDFS案例

Flume入门三_采集日志文件到HDFS

Flume 案例实操 - 实时读取本地文件到HDFS

(3)Flume监控端口,读取本地文件到HDFS,读取目录文件到HDFS

大数据实时日志收集框架Flume案例之抽取日志文件到HDFS

Flume采集目录及文件到HDFS案例